Site Reliability Engineering Team Lead (Principal SRE)
Core
Own the reliability, availability, and operational health of a global cloud-native AI platform, shaping reliability strategy across critical production services.
Role type
Principal SRE Team Lead (hands-on IC with technical leadership)
Builds
Global cloud-native AI platform services
Domain
Cloud-native AI / Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Docker, Istio, Azure, Prometheus, Grafana, Terraform, Flux, Python/Go/Shell, UNIX/Linux, high-availability architecture, CI/CD automation, SLI/SLO/SLA governance, incident response, root cause analysis
Preferred skills
Loki, Thanos, Jira, Confluence, automotive/embedded systems experience
Technologies
Kubernetes, Docker, Istio, Azure, AWS, Google Cloud, Zabbix, Prometheus, Grafana, Terraform, Flux, Python, Go, Shell, Loki, Thanos, Jira, Confluence
Responsibilities
Provide technical leadership and mentorship across distributed SRE teams; Establish technical direction, engineering standards, and reliability practices; Own and execute a 2–3 quarter reliability roadmap; Define and govern SLI/SLO/SLA frameworks for 99.95% availability; Design and maintain sustainable on-call models; Serve as Tier 2 escalation point for major production incidents; Lead blameless postmortems and drive systemic improvements; Conduct production readiness reviews; Approve high-risk production changes; Define strategic direction for metrics, dashboards, and automation; Drive CI/CD automation for deployments and rollbacks; Partner with DevOps/platform teams to evolve shared infrastructure; Incorporate reliability principles into the SDLC; Communicate technical risks and reliability posture to stakeholders.
Seniority
Principal, hands-on IC with technical leadership
