Sr Lead Site Reliability & Systems Engineer
Core
Drive platform reliability, operational excellence, and systems architecture across infrastructure to ensure scalable, resilient products.
Role type
Senior Lead Site Reliability & Systems Engineer
Builds
Highly available, fault-tolerant infrastructure on cloud platforms; CI/CD pipelines; observability platforms.
Domain
Automotive / Cloud Infrastructure / Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
SRE strategy definition, incident management, SLO/SLI management, systems architecture, infrastructure-as-code, cloud platform expertise, Python/Go/Bash/Java/C++, Kubernetes, Linux/Unix internals, networking, observability tooling, chaos engineering, team mentorship.
Preferred skills
Service mesh (Istio, Linkerd), API gateways (Kong, Apigee), systems integration, FinOps, regulated industry experience, compliance frameworks, legacy-to-cloud migrations.
Technologies
AWS, GCP, Azure, Terraform, Kubernetes, Datadog, Prometheus, Grafana, Splunk, Istio, Linkerd, Kong, Apigee
Responsibilities
Define and drive SRE strategy and roadmap; Establish and enforce SLOs, SLIs, and error budgets; Own the incident management lifecycle; Lead the design and evolution of large-scale distributed systems; Build and maintain highly available infrastructure on cloud platforms; Mentor and grow a team of SREs and platform engineers.
Seniority
Senior, hands-on IC with leadership responsibilities