Site Reliability Engineer – AI Applications (M/W/X)
Core
Ensure the reliability, scalability, and operational excellence of the AI platform powering document-rich, interactive, and intelligent use cases like hybrid search, question answering, and conversational assistants.
Role type
Senior Site Reliability Engineer (AI Infrastructure)
Builds
AI Gateway, MCP servers, LLM serving infrastructure, retrieval services, and platform components for production AI applications.
Domain
Gaming industry, AI infrastructure, distributed systems, cloud-native environments.
Deliverable
production ML models | infrastructure
Required skills
SRE practices (SLOs, error budgets, incident management), Python, Rust, Kubernetes, Docker, CI/CD (GitLab), infrastructure as code, distributed systems, microservices, networking, production debugging.
Preferred skills
AI infrastructure (AI gateways, MCP servers, LLM serving stacks), RAG and agentic architectures, GPU-backed inference, high-throughput serving, serverless systems, AI safety and governance.
Technologies
Python, Rust, GitLab CI/CD, Kubernetes, Docker, GitLab
Responsibilities
Operate and improve reliability of AI platform components; Define and maintain SLOs, SLIs, and error budgets; Build observability for AI-specific signals (latency, token throughput, costs); Improve performance and cost profile of model serving; Design graceful degradation and failover strategies; Maintain safe deployment practices; Lead incident response and capacity planning.
Seniority
Senior, hands-on IC