Staff Machine Learning Engineer – VLM/LLM Evaluation
Core
Lead the development of end-to-end evaluation systems and benchmarks for Waymo Foundation Models (LLMs/VLMs) to ensure the quality, safety, and realism of embodied AI agents used in autonomous driving.
Role type
Staff Machine Learning Engineer (Evaluation & Benchmarking)
Builds
Evaluation pipelines and benchmarks for pretraining, supervised fine-tuning, and reinforcement learning of Foundation Models
Domain
Autonomous driving / Embodied AI / Generative AI
Deliverable
production ML models
Required skills
ML engineering, applied Deep Learning, large scale distributed systems, Python, C/C++, analytical and debugging skills
Preferred skills
ML infra experience, generative models (LLMs/VLMs), reinforcement learning, ML frameworks (PyTorch, JAX, TensorFlow)
Responsibilities
Lead development of end-to-end evaluation systems and benchmarks for Waymo Foundation models; Implement and extend large scale data and evaluation pipelines; Partner across organizations to land disruptive tech in production; Work with creative teams to build state-of-the-art Foundation Models
Seniority
Staff, hands-on IC with strategic scope