Staff Engineer, Staff Engineer, Hybrid Services & Reliability (SRE)
Core
Lead reliability architecture and automation for the 'bench cloud' environment supporting autonomous vehicle testing and validation.
Role type
Staff Site Reliability Engineer (SRE)
Builds
Hybrid cloud platform for autonomous vehicle hosting, validation, and testing (Super Cruise)
Domain
Automotive / Autonomous Vehicle Infrastructure
Deliverable
infrastructure
Required skills
Site Reliability Engineering (SRE), Python, Go, Linux systems administration, DHCP/PXE automation, Infrastructure as Code, Configuration Management (Chef/Ansible)
Preferred skills
Kubernetes (k8s), hybrid cloud connectivity management
Technologies
Python, Go, Linux, DHCP, PXE, NTP, Chef, Ansible, Kubernetes
Responsibilities
Define and implement Service Level Objectives (SLOs) and SLIs for hybrid cloud services; Automate core on-prem utilities for server auto-provisioning; Develop observability dashboards and alerting to reduce Mean Time to Recovery (MTTR); Ensure data path integrity from test bench to runtime; Mentor colleagues on internal processes and services.
Seniority
Staff, hands-on IC with mentorship