Site Reliability Engineer, Inference Infrastructure
Core
Build high-performance, scalable, and reliable machine learning systems for deploying frontier AI models via API endpoints.
Role type
Site Reliability Engineer (Inference Infrastructure)
Builds
AI platform delivering large language models through API endpoints
Domain
Artificial Intelligence / Large Language Models / Cloud Infrastructure
Deliverable
production ML models
Required skills
Kubernetes, distributed systems, Linux, GCP/Azure/AWS/OCI, GPU/TPU management, Golang/C++, cost management, SLO definition
Preferred skills
Custom Kubernetes operators, multi-cloud/hybrid serving, complex Linux troubleshooting, accelerator computational characteristics
Technologies
Kubernetes, GCP, Azure, AWS, OCI, Golang, C++
Responsibilities
Build self-service systems for managing and deploying services; Automate environment observability and resilience; Participate in on-call rotation to ensure SLOs; Develop team through knowledge sharing and code reviews
Seniority
Mid-Senior, hands-on IC