Senior Site Reliability Engineer, Fleet Management
Core
Building and maintaining a scalable, secure Kubernetes runtime environment for MongoDB's internal developers to ship products.
Role type
Senior Site Reliability Engineer (Platform Engineering)
Builds
Multi-cloud Kubernetes infrastructure, networking, load balancing, and observability systems for MongoDB Atlas.
Domain
Cloud-native database infrastructure, Kubernetes, multi-cloud (AWS, GCP, Azure)
Deliverable
production ML models | infrastructure
Required skills
Kubernetes lifecycle management, Go/Python, Linux internals, networking (TCP/IP, DNS, TLS), Terraform, automation, debugging complex production issues
Preferred skills
Helm, Kustomize, Gatekeeper, Kyverno, CRDs/Operators, CRI, CSI, AWS/GCP/Azure, Crossplane, AWS Controllers for Kubernetes (ACK), namespaces, cgroups
Responsibilities
Develop and maintain scalable runtime environment on Kubernetes, provide internal support for Kubernetes ecosystem, participate in 24/7 on-call rotation, prioritize blameless post-mortems and systemic fixes
Seniority
Senior, hands-on IC