Staff Site Reliability Engineer
Core
Provide technical leadership for the operational foundations enabling deployment at scale of AI applications using MongoDB.
Role type
Staff Site Reliability Engineer (Platform Architecture & Operations)
Builds
Kubernetes fleet, networking, observability, alerting, and tenant isolation for a multi-cloud AI platform
Domain
Cloud-native data platforms, AI infrastructure, multi-cloud environments
Deliverable
production ML models | infrastructure
Required skills
Distributed systems architecture, Kubernetes platform design, Python or Go programming, workload isolation strategies, multi-cloud infrastructure reasoning, incident response leadership, SLO definition, capacity planning, automation engineering
Preferred skills
Mentoring engineering teams, cross-team collaboration, customer-focused mindset
Technologies
Kubernetes, Python, Go, AWS, GCP, Azure
Responsibilities
Own reliability architecture across regions and cloud providers, set operational standards for on-call quality and incident response, mentor the SRE team, participate in 24/7 on-call rotation, collaborate on platform operability and best practices
Seniority
Staff, hands-on IC with strategic scope