Senior Deep Learning Engineer, Accuracy Evaluation
Core
Design and build decision-grade evaluation environments and novel methodologies to assess the performance and capabilities of frontier deep learning models (LLMs, RAG, agents, vision) for NVIDIA.
Role type
Senior IC deep learning evaluation engineer
Builds
Auditable accuracy signals, benchmark environments, regression CI systems, and statistical analysis tooling for model releases.
Domain
Artificial Intelligence / Deep Learning / Model Evaluation
Deliverable
production ML models
Required skills
LLM evaluation design, statistical foundations (experimental design, significance testing, regression analysis), evaluation infrastructure development, agentic system benchmarking, low-precision inference measurement, HPC cluster management
Preferred skills
Open-source evaluation frameworks, agentic system evaluation (SWE-bench, GAIA, WebArena), evaluation research publications, quantization-aware benchmarking, MLflow/W&B experiment management
Responsibilities
Design decision-grade evaluation environments for frontier models; Research and develop novel evaluation methodologies for emerging model families; Build and operate evaluation infrastructure and pipelines; Partner with research and product teams to translate evaluation signals into release decisions.
Seniority
Senior, hands-on IC