Cloud Operations Engineer
Core
Coordinate with global teams to ensure uptime for the Atlas customer base, scale operations processes, and automate incident resolution for the MongoDB Atlas platform.
Role type
Cloud Operations Engineer (SRE/DevOps)
Builds
Monitoring systems, automation scripts, and incident response workflows for the Atlas platform
Domain
Cloud infrastructure, SRE, and database operations
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Linux system administration, networking (DNS, TCP/IP), database operations, cloud infrastructure (AWS, GCP, Azure), scripting (Java, Go, Python, Javascript), incident diagnosis and resolution
Preferred skills
MongoDB, Splunk, Kubernetes
Technologies
AWS, GCP, Azure, Linux, DNS, TCP/IP, Java, Go, Python, Javascript, MongoDB, Splunk, Kubernetes
Responsibilities
Coordinate with global teams to ensure uptime guarantees, scale operations processes, design systems to reduce Mean Time to Resolve, monitor and detect customer-facing incidents, automate routine monitoring tasks, diagnose live incidents, inform leadership of major outages, participate in on-call rotation
Seniority
Mid-level, hands-on IC