Senior Site Reliability Engineer, Vehicle SW
Core
Keep Wayve's autonomous driving fleet reliable, observable, and safe while operating on public roads by working at the boundary of software, hardware, and operations.
Role type
Senior Site Reliability Engineer (Vehicle Software)
Builds
Monitoring, logging, alerting, and on-call tooling for fleet operations, deployments, and repetitive workflows.
Domain
Autonomous driving, vehicle software, fleet operations
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Linux fundamentals, CI/CD, containers (Docker), orchestration (Kubernetes), systems/scripting languages (Python, C++, Rust), troubleshooting distributed systems and databases, observability stack design (Datadog, Prometheus, Grafana, OpenTelemetry, Splunk, Humio)
Preferred skills
Cloud platform experience (AWS, GCP, Azure), infrastructure-as-code, real-time or safety-critical systems, hardware-in-the-loop, embedded/robotics environments, fleet operations telemetry pipelines, SLO/SLI definition and reliability programs
Technologies
Docker, Kubernetes, Python, C++, Rust, Datadog, Prometheus, Grafana, OpenTelemetry, Splunk, Humio, AWS, GCP, Azure
Responsibilities
Own and improve reliability, availability, and performance of vehicle software systems; participate in on-call rotation for live systems; build and operate monitoring and alerting tooling; drive incident response and post-incident learning; design and deliver automation for fleet operations and deployments; partner with teams to define SLOs and release readiness; harden production environment through capacity planning and change management
Seniority
Senior, hands-on IC