Senior Production Engineer, Managed AI
Core
Design, operate, and scale reliable managed AI services and distributed systems for compute-intensive, latency-sensitive LLM workloads.
Role type
Senior IC production engineer (AI infrastructure)
Builds
Managed AI services, distributed AI pipelines, inference services, and observability tooling
Domain
AI infrastructure, cloud services, distributed systems
Deliverable
production ML models | product features | infrastructure
Required skills
distributed systems design, LLM/AI infrastructure experience, SRE practices (SLIs/SLOs, observability, fault tolerance), modern programming (Python/Go/Java/C++), Kubernetes/container orchestration
Preferred skills
scaling inference or training workloads for LLMs
Technologies
Kubernetes, Python, Go, Java, C++, telemetry tools
Responsibilities
Design and operate reliable managed AI services for LLM workloads; Build automation and reliability tooling for distributed AI pipelines; Define and improve SLIs/SLOs for AI workloads; Collaborate with teams to optimize large-scale training and inference clusters; Automate observability with telemetry and performance tuning; Investigate and resolve reliability issues in distributed AI systems
Seniority
Senior, hands-on IC