Site Reliability Engineer (SRE)
Core
Improve reliability, scalability, and operational excellence of a GenAI inference platform supporting the end-to-end ML lifecycle.
Role type
Site Reliability Engineer (SRE)
Builds
GenAI inference platform for deploying, scaling, and monitoring LLMs, speech, vision, and diffusion models
Domain
Generative AI / Cloud Infrastructure
Deliverable
production ML models
Required skills
Kubernetes, Linux, cloud platforms (AWS/GCP/Azure), Terraform, Helm, CI/CD pipelines, incident management, Python/Bash/Go scripting
Preferred skills
MLOps, AI infrastructure, GPU clusters, model serving systems, release engineering
Technologies
Kubernetes, AWS, GCP, Azure, Terraform, Helm, Python, Bash, Go
Responsibilities
Maintain uptime and performance of platform services; manage Kubernetes clusters and production environments; set up incident response, RCA, and deployment strategies; build monitoring and alerting dashboards; automate operational workflows; troubleshoot production issues; support GPU workloads and model serving systems
Seniority
Mid-level, hands-on IC