LEAD SITE RELIABILITY ENGINEER
Core
Lead the implementation of infrastructure automation, observability practices, and incident response workflows for enterprise scalable cloud-based solutions.
Role type
Lead Site Reliability Engineer (Manager, Non People Leader)
Builds
Enterprise scalable cloud-based solutions
Domain
Automotive / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Infrastructure as Code (Terraform, AWS CloudFormation), CI/CD tooling, observability platforms (NewRelic, Cloudwatch, Grafana, DataDog), AppSec tooling (Veracode, CloudSploit, Data Theorem), C#/Java/Python programming
Preferred skills
Unit testing, quality gates, automated functional testing, performance testing
Technologies
Terraform, AWS CloudFormation, NewRelic, Cloudwatch, Grafana, DataDog, Veracode, CloudSploit, Data Theorem, C#, Java, Python, AWS
Responsibilities
Lead implementation of IaC and automation with unit testing and quality gates; Develop and maintain observability practices including SLI/SLO/SLA and alerting; Create and manage runbooks, gamedays, postmortems, and incident response workflows; Engage with cross-functional teams to drive SRE/DevOps best practices.
Seniority
Lead, hands-on IC with management responsibilities