Staff Machine Learning Engineer - Vision-Language Foundation Models
Core
Building the cognitive engine for autonomous driving by advancing Vision-Language Models (VLM) foundation models to enable offboard reasoning, scene understanding, and data flywheel systems for the Waymo Driver.
Role type
Staff Machine Learning Engineer (Vision-Language Foundation Models)
Builds
Multimodal foundation models, offboard reasoning systems, and data flywheel pipelines for autonomous driving.
Domain
Autonomous driving, AI Foundation Models, Multimodal Learning
Deliverable
production ML models
Required skills
Foundation model lifecycle (pre-training, SFT, RL), distributed training infrastructure (FSDP, Megatron, JAX/Pax), Python/PyTorch/JAX, data curation for multimodal models, technical roadmap definition, cross-functional technical leadership
Preferred skills
PhD in CS/AI, top-tier AI publication record, advanced Reinforcement Learning paradigms, data engineering at billion/trillion token scale, multimodal perception in robotics/autonomous driving, Staff-level impact history
Technologies
Gemini, FSDP, Megatron, JAX, Pax, PyTorch, Python
Responsibilities
Lead technical strategy for massive-scale multimodal pre-training datasets and data mixture strategies; Design and implement SFT and RLHF/RLAIF/DPO/GRPO/PPO pipelines for reasoning enhancement; Architect scalable inference and evaluation pipelines for the VLM data flywheel; Define training recipes and scaling laws via ablation studies; Drive cross-functional AI strategy across ML Infra, Perception, and Behavior teams; Provide Staff-level technical leadership including roadmap ownership and mentoring senior engineers
Seniority
Staff, hands-on IC with strategic leadership