Applied AI Researcher, Agent Systems & Evaluation
Core
Build a rigorous closed-loop evaluation system for autonomous AI agents to ensure trustworthy, unattended operation within Nuro's engineering organization.
Role type
Senior Applied AI Researcher (Agent Systems & Evaluation)
Builds
Autonomous AI agent systems, evaluation pipelines, and post-trained models for internal engineering workflows.
Domain
Autonomous driving, AI agents, machine learning evaluation, and system optimization.
Deliverable
production ML models | product features | dashboards & analysis
Required skills
frontier model evaluation, closed-loop evaluation system design, automated hill climbing, supervised fine-tuning, RL on open-source VLMs, test-time scaling, inference compute optimization, statistical standards for acceptance, production trace mining, task suite construction, model-based judge validation, hypothesis testing against production traffic, reading and replicating research frontier.
Preferred skills
engineering background, strong research taste, experience with autonomous driving data, labeling workforce management.
Technologies
open-source vision-language models, proprietary driving data, RL, supervised fine-tuning, VLMs.
Responsibilities
Own the evaluation pipeline end-to-end (data collection, loop construction, automated hill climbing), establish evaluation foundations (data sources, task suites, noise floors), run post-training experiments on VLMs, convert research frontier into live experiments, optimize inference compute and sampling strategies, quantify platform impact with confidence intervals.
Seniority
Senior, hands-on IC with strategic impact