CareerPlanSign in

Site Reliability Engineer (SRE)

💼 Full-time🗓 2026-09-28

Core

Improve reliability, scalability, and operational excellence of a GenAI inference platform supporting the end-to-end ML lifecycle.

Role type

Site Reliability Engineer (SRE)

Builds

GenAI inference platform for deploying, scaling, and monitoring LLMs, speech, vision, and diffusion models

Domain

Generative AI / Cloud Infrastructure

Deliverable

production ML models

Required skills

Kubernetes, Linux, cloud platforms (AWS/GCP/Azure), Terraform, Helm, CI/CD pipelines, incident management, Python/Bash/Go scripting

Preferred skills

MLOps, AI infrastructure, GPU clusters, model serving systems, release engineering

Technologies

Kubernetes, AWS, GCP, Azure, Terraform, Helm, Python, Bash, Go

Responsibilities

Maintain uptime and performance of platform services; manage Kubernetes clusters and production environments; set up incident response, RCA, and deployment strategies; build monitoring and alerting dashboards; automate operational workflows; troubleshoot production issues; support GPU workloads and model serving systems

Seniority

Mid-level, hands-on IC

Sourced via wellfound · Listed on CareerPlan, which tracks 848,000+ jobs from 20+ sources.