Site Reliability Engineer - Career
Core
Manage system uptime across cloud-native and hybrid architectures, build infrastructure as code, and lead incident response for distributed services.
Role type
Senior Site Reliability Engineer (IC)
Builds
Scalable cloud infrastructure, CI/CD pipelines, automated deployment tooling, and comprehensive runbooks.
Domain
Cloud-native infrastructure, DevOps, and distributed systems
Deliverable
production ML models | product features | infrastructure
Required skills
Infrastructure as Code (Terraform), CI/CD pipeline design, Linux/Windows system administration, container orchestration (Kubernetes), scripting (Python, Bash, Go, Java), monitoring and observability, incident triage, blameless postmortems
Preferred skills
Cloud certifications (CKA, CKAD), experience in highly secure/regulated industries, deep expertise in large-scale distributed systems, mentorship capabilities
Technologies
AWS, GCP, Terraform, Jenkins, Docker, Kubernetes, Python, Bash, Go, Java, Node.js, Chef, Ansible
Responsibilities
Manage system uptime across cloud-native and hybrid architectures; Build infrastructure as code patterns meeting security standards; Build CI/CD pipelines for application and cloud architecture; Solve problems and triage complex distributed architecture service maps; Lead availability blameless postmortems and drive remediation
Seniority
Senior, hands-on IC