Principal Software Engineer - DevOps / Site Reliability Engineer
Core
Own and evolve operational foundations for the AI Efficiency team's tech platform and tools to ensure reliability, scalability, and security at growing scale.
Role type
Principal DevOps / Site Reliability Engineer
Builds
Production AI services, internal tools, and the underlying infrastructure supporting them
Domain
Gaming industry, Cloud Infrastructure, AI Engineering
Deliverable
production ML models | infrastructure
Required skills
Cloud platform design (AWS/GCP/Azure), CI/CD pipeline design, Observability (metrics/logs/traces), Infrastructure as Code, Capacity planning, Incident management, Automation scripting (Python/Go/JS/TS), Disaster recovery design, Service resilience patterns
Preferred skills
AI-native workflow operationalization, Agentic system guardrails, Progressive delivery strategies, Self-service platform development
Technologies
Python, Go, JavaScript, TypeScript, AWS, GCP, Azure, Kubernetes, Docker, Terraform, Prometheus, Grafana, ELK, Jenkins, GitHub Actions
Responsibilities
Design and maintain infrastructure for production AI services; Improve CI/CD pipelines and deployment automation; Define and operationalize SLOs, SLIs, and error budgets; Build comprehensive observability using metrics, logs, and traces; Lead incident response and post-mortem reviews; Automate operational toil and reduce manual intervention; Design resilience and disaster recovery strategies; Improve developer experience via self-service tooling; Partner with ML engineers on model-serving reliability; Implement AI-assisted operational workflows with proper guardrails; Mentor team members on operational excellence.
Seniority
Principal, hands-on IC with strategic leadership