Site Reliability Engineer (Senior or Staff), Storage Layer Services (SLS)
Core
Re-architecting MongoDB's cloud storage layer to build performant, multi-tenant distributed storage services that enhance the Atlas stack and enable efficient customer workloads.
Role type
Senior/Staff Site Reliability Engineer (Storage Layer Services)
Builds
Multi-tenant distributed storage services for MongoDB Atlas
Domain
Cloud infrastructure, distributed systems, database storage
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
distributed systems operations, Python/Go programming, stateful storage systems, Kubernetes, cloud platforms (AWS/GCP/Azure), Linux internals, networking (TCP/IP/DNS/TLS)
Preferred skills
architectural shifts, multi-cloud infrastructure management, secure multi-tenant runtime design
Responsibilities
Define SLOs and capacity plans, ensure reliability and durability of storage layer, identify metrics for incident detection, participate in 24/7 on-call rotation, optimize infrastructure performance from application to kernel
Seniority
Senior/Staff, hands-on IC