Site Reliability Engineering Team Lead (Principal SRE)
Core
Lead the Site Reliability Engineering team to own the reliability, availability, and operational health of Cerence's cloud-native automotive AI platform (voice, gesture, gaze solutions).
Role type
Principal SRE Team Lead (hands-on IC with technical authority)
Builds
Cloud-native automotive AI systems for global automakers
Domain
Automotive AI / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Site reliability engineering, cloud platform management, container orchestration (Kubernetes, Docker, Istio), public cloud expertise (Azure, AWS, GCP), observability tooling (Prometheus, Grafana, Zabbix), CI/CD automation, infrastructure-as-code (Terraform, Flux), scripting/programming (Python, Go, Shell), UNIX/Linux system administration
Preferred skills
Distributed team leadership, high-availability service design, log aggregation (Loki, Thanos), automotive/embedded domain experience
Technologies
Kubernetes, Docker, Istio, Azure, AWS, Google Cloud, Prometheus, Grafana, Zabbix, Terraform, Flux, Python, Go, Shell, Loki, Thanos
Responsibilities
Define and govern SLI/SLO/SLA frameworks; serve as Tier 2 technical escalation point for major incidents; drive reliability roadmap and production readiness reviews; set strategic direction for observability and automation; mentor and technically develop the team; approve high-risk production changes
Seniority
Principal, hands-on IC
