Senior Research Engineer - Enterprise Products
Core
Designing optimized inference technologies and routing policies for Large Language Models (LLMs) to support NVIDIA's generative AI products.
Role type
Senior Research Engineer (Generative AI Inference)
Builds
Optimized inference software, routing policies, agentic benchmarks, and open-source contributions.
Domain
Generative AI, Deep Learning, Natural Language Processing
Deliverable
production ML models
Required skills
Deep Learning frameworks (PyTorch, TensorFlow), LLM evaluation/benchmarking, statistical analysis, Machine Learning, Deep Neural Networks, Natural Language Processing, empirical research, algorithms and data structures, parallel and distributed computing, system software
Preferred skills
Large-scale distributed systems architecture for deep learning, agentic benchmark creation and publications, CPU/GPU architecture knowledge, GPU programming (CUDA), mentoring
Technologies
PyTorch, TensorFlow, CUDA, NVIDIA accelerated serving stack
Responsibilities
Design and evaluate routing policies for LLM traffic to best use mixture of model systems; Build and run agentic benchmarks to measure algorithm quality and generate calibration data; Ship design docs, code, and community contributions to open-source repos; Collaborate with engineering teams to ensure software integration across the NVIDIA stack
Seniority
Senior, hands-on IC