Staff Software Engineer, Compute Reliability
Core
Design and develop instrumentation and onboard architecture to prevent, detect, and debug reliability issues across Waymo compute platforms.
Role type
Staff Software Engineer, Compute Reliability
Builds
Scalable tools to automate triage, debugging, and resolution of compute hardware reliability and numerical stability issues.
Domain
Autonomous driving / High-performance computing
Deliverable
production ML models | infrastructure
Required skills
C++, CUDA, GPU/TPU acceleration, GPU kernel development, debugging GPU workloads, compiler optimization (LLVM, XLA, JAX), memory management, concurrency control
Preferred skills
Autonomous vehicles (L4) experience, ADAS systems (L2/L3) experience, software reliability expertise
Technologies
CUDA, TPU, GPU, LLVM, XLA, JAX, AutoFDO
Responsibilities
Drive design and development of instrumentation and onboard architecture; Design and implement long-term strategies to mitigate compute hardware reliability and numerical stability issues; Lead cross-functional initiatives to architect proactive solutions for compiler and optimization issues; Architect frameworks to identify and eliminate complex software and firmware faults; Influence the design and tooling of next-generation hardware platforms
Seniority
Staff, hands-on IC with strategic influence