Senior Manager - Reliability Engineering (Member Experience)
Core
Lead the Site Reliability Engineering function to set reliability standards and support the streaming architecture for Netflix's global platform.
Role type
Senior Manager, Site Reliability Engineering
Builds
Streaming media infrastructure and reliability operating models
Domain
Streaming media, Cloud-native infrastructure
Deliverable
production ML models | infrastructure
Required skills
Cloud-native scale (AWS/GCP, Containers, Service Mesh), Modern observability (Metrics, Tracing, Logging), Strategic leadership, Reliability governance (SLIs/SLOs, Error Budgets), Cross-functional partnership, Organizational influence
Preferred skills
Streaming media infrastructure, LLM-based autonomous agents, Netflix OSS ecosystem (Spinnaker, Atlas, Mantis, Chaos Monkey), SRE AI fluency
Technologies
AWS, GCP, Containers, Service Mesh, Spinnaker, Atlas, Mantis, Chaos Monkey
Responsibilities
Build and scale a world-class SRE function, Establish company-wide reliability standards and scorecards, Collaborate with CDN, Playback, and Ads teams to eliminate systemic failures
Seniority
Senior, hands-on IC with leadership mandate