Site Reliability Engineers (SRE)
Core
Drive cloud resiliency through observability, chaos engineering, and migration of workloads to AWS.
Role type
Senior Site Reliability Engineer (Cloud Resiliency & Observability)
Builds
Cloud infrastructure, observability platforms, and resilient systems for the organization
Domain
Cloud computing, Observability, Chaos Engineering
Deliverable
production ML models | product features | infrastructure
Required skills
Chaos Engineering, Cloud migration, Observability design, Kubernetes, Infrastructure as Code (Terraform, CloudFormation), Linux, Networking, Scripting, Incident management, Mentoring
Preferred skills
Kafka (MSK), RDBMS (Postgres, MySQL), Python, Go
Technologies
AWS, Kubernetes, Terraform, CloudFormation, Git, Kafka, Postgres, MySQL, Python, Go
Responsibilities
Conduct Chaos Engineering experiments, Research and monitor cloud migrations, Design and refactor observability, Coordinate observability across teams, Assess cloud deployments for compliance, Investigate and correct observability gaps, Participate in 24/7 on-call rotation, Mentor colleagues on technical skills
Seniority
Senior, hands-on IC