CareerPlanSign in

Senior Production Engineer, Managed Cloud

San Francisco, CA - US💼 Full-time💰 $170,000–$205,000🗓 2026-07-29 → 2026-09-27

Core

Design and operate reliable managed AI services to serve and scale LLM workloads, ensuring high availability and performance for compute-intensive, latency-sensitive AI infrastructure.

Role type

Senior Production Engineer (AI Infrastructure)

Builds

Managed AI services, distributed training and inference clusters, observability systems

Domain

AI Infrastructure, Cloud Services, Distributed Systems

Deliverable

production ML models | infrastructure

Required skills

Distributed systems design, SRE practices (SLIs/SLOs, fault tolerance), observability/telemetry, modern programming (Python/Go/Java/C++), Kubernetes/container orchestration, production-grade system development

Preferred skills

None explicitly stated

Technologies

Kubernetes, Python, Go, Java, C++, LLMs

Responsibilities

Design and operate reliable managed AI services, define and measure SLIs/SLOs, optimize large-scale training and inference clusters, build telemetry and performance tuning strategies, investigate and resolve reliability issues in distributed AI systems, contribute to next-generation distributed system architecture

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.