Senior Site Reliability Engineer, Fleet Management
Core
Building and maintaining a scalable, secure Kubernetes runtime environment for MongoDB's product teams to support distributed systems and AI-era applications.
Role type
Senior Site Reliability Engineer (Platform Engineering)
Builds
Multi-cloud Kubernetes infrastructure, networking, load balancing, and observability systems for MongoDB Atlas.
Domain
Cloud-native infrastructure, Kubernetes, multi-cloud (AWS, GCP, Azure), and distributed systems.
Deliverable
production ML models | product features | infrastructure
Required skills
Kubernetes lifecycle management, Go/Python programming, Linux internals, networking (TCP/IP, DNS, TLS), containerization, Infrastructure as Code (Terraform), automation, debugging complex production issues, blameless post-mortems.
Preferred skills
Kubernetes Operators (Helm, Kustomize, Gatekeeper, Kyverno), CRDs, CRI, CSI, multi-tenant runtime design, advanced Linux (namespaces, cgroups), cloud platforms (AWS, GCP, Azure), Crossplane, AWS Controllers for Kubernetes (ACK).
Responsibilities
Develop and maintain scalable Kubernetes runtime environments; provide internal support for the Kubernetes ecosystem; participate in 24/7 on-call rotation to resolve critical issues; prioritize blameless post-mortems and systemic fixes.
Seniority
Senior, hands-on IC