Site Reliability Engineer (Senior or Staff), Atlas
Core
Design, build, and maintain a reliable, resilient multi-cloud database platform (Atlas) serving business-critical workloads for a wide range of customers.
Role type
Senior Site Reliability Engineer (Infrastructure)
Builds
A globally distributed, multi-cloud database platform available on AWS, Google Cloud, and Microsoft Azure.
Domain
Cloud-native database infrastructure, multi-cloud environments.
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Linux system administration, cloud platform expertise (AWS/Azure/GCP), modern programming (Go/Ruby/Python), web and network protocols (HTTP/TLS/DNS), automation, on-call incident response.
Preferred skills
None stated.
Technologies
AWS, Azure, GCP, Linux, Go, Ruby, Python, HTTP, TLS, DNS.
Responsibilities
Design and build complex systems for the Atlas fleet, collaborate with software engineering teams to solve technical challenges, participate in 24/7 on-call rotation to resolve disruptions.
Seniority
Senior, hands-on IC