Senior Deep Learning Researcher, LLM Inference
Core
Conducting research and developing next-generation algorithms to optimize large language model (LLM) inference for low-latency and high-throughput scenarios.
Role type
Senior Deep Learning Researcher (LLM Inference)
Builds
NVIDIA software products and customer-facing inference solutions
Domain
Artificial Intelligence / Large Language Models / High-Performance Computing
Deliverable
production ML models
Required skills
LLM architecture design, deep learning research, Python programming, PyTorch, algorithmic optimization, software engineering standards, HPC environments, large-scale GPU cluster management, LLM inference systems (vLLM, TensorRT-LLM)
Preferred skills
Speculative decoding, parallelization strategies, world-class industrial research group experience, top-tier institution experience
Technologies
PyTorch, vLLM, TensorRT-LLM, Python
Responsibilities
Research and implement groundbreaking algorithms for LLM inference, translate research into practical software solutions, analyze algorithm performance on NVIDIA hardware to identify bottlenecks, collaborate with internal research and engineering teams, partner with scientific organizations to integrate innovations
Seniority
Senior, hands-on IC