CareerPlanGet AI match score →

Cloud Site Reliability Engineer

San Jose, CA💼 Full-time🗓 2026-07-08 → 2026-07-31

Core

Guardian of reliability, performance, and scalability for a full-stack generative AI inferencing service, ensuring exceptional uptime and low-latency response times for enterprise and government customers.

Role type

Cloud Site Reliability Engineer (SRE) specializing in AI inferencing

Builds

Production AI inferencing endpoints and the underlying cloud infrastructure for a full-stack generative AI platform

Domain

Generative AI / High-Performance Computing / Cloud Infrastructure

Deliverable

production ML models

Required skills

Python, Go, or Java; Kubernetes; Prometheus; Terraform; CI/CD; distributed systems troubleshooting

Preferred skills

Hybrid cloud/on-prem infrastructure; ML/AI inferencing support; GPU-accelerated computing; vLLM, SGLang, or Ray; MLOps; SQL/NoSQL databases; Redis/Memcached caching

Technologies

AWS, GCP, Azure, Docker, Kubernetes, Prometheus, Grafana, Datadog, Terraform, Ansible, Jenkins, GitHub Actions, ArgoCD, vLLM, SGLang, Ray, Redis, Memcached

Responsibilities

Manage production inferencing service availability, latency, and performance across multiple regions; Lead incident response and drive blameless post-mortems; Develop and maintain advanced monitoring, alerting, and dashboarding; Design and implement auto-scaling policies to handle variable inference loads; Manage and evolve cloud infrastructure using Infrastructure as Code; Build and improve CI/CD pipelines for model version deployment; Forecast infrastructure needs and optimize cloud costs; Define, measure, and report on Service Level Objectives (SLOs) and Indicators (SLIs)

Seniority

Mid-level, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗