Cloud Site Reliability Engineer
Core
Guardian of reliability, performance, and scalability for a full-stack generative AI inferencing service, ensuring exceptional uptime and low-latency response times for enterprise and government customers.
Role type
Cloud Site Reliability Engineer (SRE) specializing in AI inferencing
Builds
Production AI inferencing endpoints and the underlying cloud infrastructure for a full-stack generative AI platform
Domain
Generative AI / High-Performance Computing / Cloud Infrastructure
Deliverable
production ML models
Required skills
Python, Go, or Java; Kubernetes; Prometheus; Terraform; CI/CD; distributed systems troubleshooting
Preferred skills
Hybrid cloud/on-prem infrastructure; ML/AI inferencing support; GPU-accelerated computing; vLLM, SGLang, or Ray; MLOps; SQL/NoSQL databases; Redis/Memcached caching
Technologies
AWS, GCP, Azure, Docker, Kubernetes, Prometheus, Grafana, Datadog, Terraform, Ansible, Jenkins, GitHub Actions, ArgoCD, vLLM, SGLang, Ray, Redis, Memcached
Responsibilities
Manage production inferencing service availability, latency, and performance across multiple regions; Lead incident response and drive blameless post-mortems; Develop and maintain advanced monitoring, alerting, and dashboarding; Design and implement auto-scaling policies to handle variable inference loads; Manage and evolve cloud infrastructure using Infrastructure as Code; Build and improve CI/CD pipelines for model version deployment; Forecast infrastructure needs and optimize cloud costs; Define, measure, and report on Service Level Objectives (SLOs) and Indicators (SLIs)
Seniority
Mid-level, hands-on IC