Senior Site Reliability Engineer, Vehicle SW
Core
Keep Wayve's autonomous driving fleet reliable, observable, and safe while operating on public roads by working at the boundary of software, hardware, and operations.
Role type
Senior Site Reliability Engineer (Vehicle Software)
Builds
Monitoring, logging, alerting, and on-call tooling for fleet operations and deployments
Domain
Autonomous driving / Embodied AI / Vehicle Software
Deliverable
production ML models | infrastructure
Required skills
Linux fundamentals, CI/CD, containers (Docker), orchestration (Kubernetes), systems/scripting languages (Python, C++, Rust), deep troubleshooting (networking, distributed systems, databases), observability stack design (Datadog, Prometheus, Grafana, OpenTelemetry, Splunk, Humio)
Preferred skills
Cloud platform experience (AWS, GCP, Azure) with IaC, real-time/safety-critical systems, hardware-in-the-loop, embedded/robotics environments, fleet operations/telemetry pipelines, SLO/SLI definition
Technologies
Docker, Kubernetes, Python, C++, Rust, Datadog, Prometheus, Grafana, OpenTelemetry, Splunk, Humio, AWS, GCP, Azure
Responsibilities
Own and improve reliability, availability, and performance of vehicle software systems; participate in on-call rotation for live systems; build and operate monitoring, logging, and alerting tooling; drive incident response and post-incident learning; design and deliver automation for fleet operations and deployments; partner with teams to define SLOs and release readiness; harden production environment through capacity planning and change management
Seniority
Senior, hands-on IC