CareerPlanSign in

Site Reliability Engineering Team Lead (Principal SRE)

US🌐 Remote💼 Full-time🗓 2026-09-18 → 2026-09-25

Core

Own the reliability, availability, and operational health of a global cloud-native AI platform, shaping reliability strategy across critical production services.

Role type

Principal SRE Team Lead (hands-on IC with technical leadership)

Builds

Global cloud-native AI platform services

Domain

Cloud-native AI / Distributed Systems

Deliverable

production ML models | infrastructure

Required skills

Kubernetes, Docker, Istio, Azure, Prometheus, Grafana, Terraform, Flux, Python/Go/Shell, UNIX/Linux, high-availability architecture, CI/CD automation, SLI/SLO/SLA governance, incident response, root cause analysis

Preferred skills

Loki, Thanos, Jira, Confluence, automotive/embedded systems experience

Technologies

Kubernetes, Docker, Istio, Azure, AWS, Google Cloud, Zabbix, Prometheus, Grafana, Terraform, Flux, Python, Go, Shell, Loki, Thanos, Jira, Confluence

Responsibilities

Provide technical leadership and mentorship across distributed SRE teams; Establish technical direction, engineering standards, and reliability practices; Own and execute a 2–3 quarter reliability roadmap; Define and govern SLI/SLO/SLA frameworks for 99.95% availability; Design and maintain sustainable on-call models; Serve as Tier 2 escalation point for major production incidents; Lead blameless postmortems and drive systemic improvements; Conduct production readiness reviews; Approve high-risk production changes; Define strategic direction for metrics, dashboards, and automation; Drive CI/CD automation for deployments and rollbacks; Partner with DevOps/platform teams to evolve shared infrastructure; Incorporate reliability principles into the SDLC; Communicate technical risks and reliability posture to stakeholders.

Seniority

Principal, hands-on IC with technical leadership

Sourced via lever · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.