Staff ML Engineer, Generative Model Performance & Efficiency
Core
Optimizing the performance, scalability, and efficiency of generative ML models (Transformers, Diffusion, MoEs) used in Waymo's autonomous driving simulation environments.
Role type
Staff ML Engineer (Generative Model Performance & Efficiency)
Builds
High-fidelity 3D virtual worlds and simulation environments for testing and validating the Waymo Driver.
Domain
Autonomous driving / Generative AI / Simulation
Deliverable
production ML models
Required skills
Deep learning architectures (Transformers, Diffusion, MoEs), model optimization (quantization, pruning, distillation), hardware acceleration (TPU, GPU), distributed training strategies, Python, C++, profiling tools (XProf, Perfetto, Nsight)
Preferred skills
ML compilers (XLA), multi-device training/serving, framework contributions
Technologies
JAX, Flax, TensorFlow, PyTorch, XLA, TPU, GPU
Responsibilities
Analyze model architectures to identify training/inference bottlenecks; Develop techniques for model compression and efficiency; Optimize code for specific hardware accelerators; Design model partitioning and sharding strategies; Implement low-latency serving solutions; Build performance analysis and debugging tools
Seniority
Staff, hands-on IC