Senior Machine Learning Engineer, Runtime and Serving
Core
Architect and develop high-performance ML runtime and serving systems for Waymo's autonomous driving models, optimizing for both onboard vehicle compute and offboard data center environments.
Role type
Senior IC machine learning systems engineer (runtime & serving)
Builds
Efficient ML inference engines and serving infrastructure for autonomous vehicle workloads
Domain
Autonomous driving / Deep learning systems
Deliverable
production ML models
Required skills
C++ (5+ years), Python, deep learning frameworks (PyTorch, JAX), ML compiler/runtime modification, hardware accelerator optimization (GPUs, TPUs), low-latency distributed systems, profiling and benchmarking
Preferred skills
PhD in CS/EE/Deep Learning, LLM serving system experience, custom kernel development (CUDA, Triton, JAX/Pallas), unified serving API architecture
Technologies
JAX, XLA, Triton, CUDA, OpenXLA, PjRT, TensorRT, PyTorch
Responsibilities
Architect efficient ML runtime systems for onboard and offboard environments; Lead integration of ML inference runtimes; Drive migration to JAX-native runtime architecture; Collaborate on hardware-aware compute optimizations; Design tooling for profiling and bottleneck identification
Seniority
Senior, hands-on IC