Senior Site Reliability Engineer
Core
Build and operate the operational foundations (Kubernetes fleet, networking, observability) for a new platform enabling customers to build AI applications using MongoDB.
Role type
Senior Site Reliability Engineer (Infrastructure)
Builds
Multi-tenant Kubernetes infrastructure running customer AI workloads
Domain
Cloud-native data platform for AI applications
Deliverable
production ML models | infrastructure
Required skills
distributed systems operations, Kubernetes production operations, cloud infrastructure (AWS/GCP/Azure), Linux internals, TCP/IP/DNS/TLS/routing, Python/Go
Preferred skills
Kubernetes networking (Istio/Cilium), service mesh, edge load balancing, secure multi-tenant runtime environments, multi-cloud management, workload isolation
Technologies
Kubernetes, AWS, Google Cloud Platform, Azure, Python, Go, Istio, Cilium
Responsibilities
Operate and improve multi-tenant Kubernetes infrastructure, build resilient and self-healing services, identify and configure key metrics for incident detection, participate in 24/7 on-call rotation, mentor early-career SREs
Seniority
Senior, hands-on IC