Site Reliability Engineer - Singapore
Core
Ensure reliability and performance of AI products at scale, lead incident response, and build automation for system resilience.
Role type
Senior Site Reliability Engineer
Builds
AI products and infrastructure
Domain
AI/ML, Cloud Infrastructure, Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, distributed systems, cloud platforms (AWS/GCP/Azure), incident management, automation scripting, programming (Go/Python/Java)
Preferred skills
AI/ML platform support, SLO/SLA frameworks, multi-region systems
Technologies
Kubernetes, AWS, GCP, Azure, Go, Python, Java
Responsibilities
Ensure reliability and performance of AI products at scale, lead incident response and continuous reliability improvement, build automation to reduce toil, partner with product and engineering teams on reliability design, improve observability and operational maturity
Seniority
Senior, hands-on IC