CareerPlanGet AI match score →

Senior Research Engineer - Enterprise Products

3 Locations💼 Full-time💰 $192,000–$192,000🗓 2026-07-10 → 2026-08-01

Core

Designing optimized inference technologies and routing policies for Large Language Models (LLMs) to support NVIDIA's generative AI products.

Role type

Senior Research Engineer (Generative AI Inference)

Builds

Optimized inference software, routing policies, agentic benchmarks, and open-source contributions.

Domain

Generative AI, Deep Learning, Natural Language Processing

Deliverable

production ML models

Required skills

Deep Learning frameworks (PyTorch, TensorFlow), LLM evaluation/benchmarking, statistical analysis, Machine Learning, Deep Neural Networks, Natural Language Processing, empirical research, algorithms and data structures, parallel and distributed computing, system software

Preferred skills

Large-scale distributed systems architecture for deep learning, agentic benchmark creation and publications, CPU/GPU architecture knowledge, GPU programming (CUDA), mentoring

Technologies

PyTorch, TensorFlow, CUDA, NVIDIA accelerated serving stack

Responsibilities

Design and evaluate routing policies for LLM traffic to best use mixture of model systems; Build and run agentic benchmarks to measure algorithm quality and generate calibration data; Ship design docs, code, and community contributions to open-source repos; Collaborate with engineering teams to ensure software integration across the NVIDIA stack

Seniority

Senior, hands-on IC

Rewrite
## Responsibilities - Design and evaluate routing policies for LLM traffic to best use mixture of model systems. - Build and run agentic benchmarks (e.g., Terminal-Bench) to measure algorithm quality, and turn results into calibration data and routing profiles - Ship to an open-source repo: design docs, code review, docs, and community contributions - Collaborating with engineering teams across all of NVIDIA to ensure our software integrates seamlessly up and down the NVIDIA accelerated serving stack. ## Requirements - Bachelor's of Master's degree in Computer Science or equivalent experience. - 8+ years of industry experience in Deep Learning frameworks (PyTorch or TensorFlow). - Experience designing or running LLM evaluations/benchmarks — ideally agentic ones — and drawing statistically sound conclusions from them - Understanding of modern techniques in Machine Learning, Deep Neural Networks, Natural Language Processing, or Speech Recognition. - Empirical research mindset: forming hypotheses about new algorithms, running calibrations, iterating on results - Strong communication and interpersonal skills, along with the ability to work in a dynamic and distributed team. A history of mentoring junior engineers and interns is a huge plus. - A desire to constantly grow and learn new things. - Strong computer science fundamentals - algorithms and data structures, computational complexity, parallel and distributed computing, system software. ## Nice to Have - Experience architecting or developing large-scale distributed systems for deep learning. - Agentic benchmark creation and publications. - Knowledge of CPU and/or GPU architecture. - GPU programming (CUDA). ## Benefits - Base salary range: 192,000 USD - 304,750 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. - Eligibility for equity and benefits. - Applications accepted until July 14, 2026. - NVIDIA uses AI tools in its recruiting processes. - NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗