Site Reliability Engineer (Senior or Staff), Storage Layer Services (SLS)
Core
Re-architecting MongoDB's cloud storage layer to build performant, multi-tenant distributed storage services that enhance the Atlas stack and enable efficient customer workloads.
Role type
Senior/Staff Site Reliability Engineer (Storage Layer)
Builds
Multi-tenant distributed storage services for MongoDB Atlas
Domain
Cloud Infrastructure / Distributed Systems / Database Storage
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Distributed systems operations, Python/Go, Stateful storage systems, Kubernetes, Cloud platforms (AWS/GCP/Azure), Linux internals, Networking (TCP/IP/DNS/TLS)
Preferred skills
Architectural shifts, Multi-cloud management, Secure multi-tenant runtime design
Responsibilities
Define SLOs and capacity plans, Ensure reliability and durability of storage layer, Identify metrics for incident detection, Participate in 24/7 on-call rotation, Optimize infrastructure performance from application to kernel
Seniority
Senior/Staff, hands-on IC