Senior Research Manager, World Model Evaluation
Core
Lead world-model evaluation and benchmarking for NVIDIA's Physical AI portfolio, defining scientific roadmaps, discovering model failures, and driving improvement loops between evaluation and training.
Role type
Senior Research Manager (Physical AI / World Models)
Builds
Closed and open-system evaluation benchmarks, diagnostic tools, failure taxonomies, and evaluation-to-training feedback loops for world foundation models and robotics.
Domain
Physical AI, Multimodal AI, Robotics, World Foundation Models
Deliverable
production ML models
Required skills
Machine learning, computer vision, multimodal AI, robotics, world models, representation learning, model evaluation, mechanistic interpretability, team leadership, scientific roadmap definition, benchmark design, causal analysis, data curation
Preferred skills
Experience with video models, vision-language-action models, diffusion/flow models, self-supervised learning, open-system analysis techniques, publication record in top venues
Technologies
Sparse autoencoders, attention analysis, causal interventions, activation patching, representation probing, simulation systems
Responsibilities
Lead a team of Research Scientists in world-model evaluation and diagnostics; Define scientific roadmaps for closed/open-system benchmarks and metrics; Develop benchmarks for physical plausibility, temporal consistency, and spatial reasoning; Develop mechanistic evaluation methods using model internals; Drive evaluation-to-model-improvement loops with training and data teams; Publish papers, technical reports, and open-source evaluation artifacts.
Seniority
Senior, hands-on IC with management responsibilities