Site Reliability Engineer
Core
Build and maintain automation for a managed distributed SQL database service running on Kubernetes across AWS, Azure, and Google Cloud.
Role type
Senior Site Reliability Engineer (Cloud Infrastructure)
Builds
Automated infrastructure rollouts, telemetry platforms, and cloud-native database services
Domain
Cloud Infrastructure / Distributed Databases
Deliverable
production ML models | product features | infrastructure
Required skills
Infrastructure automation, Python, Golang, Kubernetes, Cloud platforms (AWS/Azure/GCP), Production debugging, Postmortem analysis
Preferred skills
Deep understanding of query engines, Backend infrastructure
Technologies
Kubernetes, Python, Golang, AWS, Azure, Google Cloud
Responsibilities
Develop automation platforms for infrastructure rollouts, Optimize telemetry platforms for event identification, Partner with engineering to optimize cloud service performance, Debug live site events and conduct RCA analysis, Participate in SLA-driven on-call rotation
