Site Reliability Engineer
Core
Build and maintain a monitorable, performant, reliable, and highly-scalable deployment platform for thousands of nodes and pods across multiple clusters to support live rides and special large events.
Role type
Site Reliability Engineer (Infrastructure & Kubernetes)
Builds
Critical infrastructure for tens of thousands of pods, auto-scaling systems for live events, and a platform for machine learning and other workloads.
Domain
Cloud Infrastructure / DevOps / Streaming
Deliverable
infrastructure
Required skills
Kubernetes cluster management, observability and monitoring at scale, Kubernetes security, CI/CD systems, Infrastructure as Code, Python or Golang, root-cause analysis, system design for reliability and capacity.
Preferred skills
Experience transitioning development teams to container-native environments, automation of day-to-day tasks, operational security and compliance expertise.
Technologies
Kubernetes, Jenkins, ArgoCD, Harness, Tekton, Terraform, Pulumi, Python, Golang, Java, C, Ubuntu, Nginx, Chef, Amazon Web Services.
Responsibilities
Host critical infrastructure ensuring best member experience on tens of thousands of pods; provide a platform for machine learning and other workloads; promote best practices for building and operating highly reliable systems; serve as domain expert in observability and monitoring; consult in system design to meet reliability and capacity requirements; automate everything from infrastructure down to day-to-day tasks; conduct timely post-mortems of infrastructure incidents; assist with all aspects of operational security and compliance; seek out potential threats to security and reliability and advocate solutions.
Seniority
Mid-to-Senior, hands-on IC