CareerPlanSign in

Staff Machine Learning Engineer - Vision-Language Foundation Models

Mountain View (US-MTV-EMF680)💼 Full-time💰 $251,000–$251,000🗓 2026-07-23 → 2026-09-26

Core

Building the cognitive engine for autonomous driving by advancing Vision-Language Models (VLM) foundation models to enable offboard reasoning, scene understanding, and data flywheel systems for the Waymo Driver.

Role type

Staff Machine Learning Engineer (Vision-Language Foundation Models)

Builds

Multimodal foundation models, offboard reasoning systems, and data flywheel pipelines for autonomous driving.

Domain

Autonomous driving, AI Foundation Models, Multimodal Learning

Deliverable

production ML models

Required skills

Foundation model lifecycle (pre-training, SFT, RL), distributed training infrastructure (FSDP, Megatron, JAX/Pax), Python/PyTorch/JAX, data curation for multimodal models, technical roadmap definition, cross-functional technical leadership

Preferred skills

PhD in CS/AI, top-tier AI publication record, advanced Reinforcement Learning paradigms, data engineering at billion/trillion token scale, multimodal perception in robotics/autonomous driving, Staff-level impact history

Technologies

Gemini, FSDP, Megatron, JAX, Pax, PyTorch, Python

Responsibilities

Lead technical strategy for massive-scale multimodal pre-training datasets and data mixture strategies; Design and implement SFT and RLHF/RLAIF/DPO/GRPO/PPO pipelines for reasoning enhancement; Architect scalable inference and evaluation pipelines for the VLM data flywheel; Define training recipes and scaling laws via ablation studies; Drive cross-functional AI strategy across ML Infra, Perception, and Behavior teams; Provide Staff-level technical leadership including roadmap ownership and mentoring senior engineers

Seniority

Staff, hands-on IC with strategic leadership

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.