Senior Production Engineer, Managed Cloud
Core
Design and operate reliable managed AI services to serve and scale LLM workloads, ensuring high availability and performance for compute-intensive, latency-sensitive AI infrastructure.
Role type
Senior Production Engineer (AI Infrastructure)
Builds
Managed AI services, distributed training and inference clusters, observability systems
Domain
AI Infrastructure, Cloud Services, Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Distributed systems design, SRE practices (SLIs/SLOs, fault tolerance), observability/telemetry, modern programming (Python/Go/Java/C++), Kubernetes/container orchestration, production-grade system development
Preferred skills
None explicitly stated
Technologies
Kubernetes, Python, Go, Java, C++, LLMs
Responsibilities
Design and operate reliable managed AI services, define and measure SLIs/SLOs, optimize large-scale training and inference clusters, build telemetry and performance tuning strategies, investigate and resolve reliability issues in distributed AI systems, contribute to next-generation distributed system architecture
Seniority
Senior, hands-on IC