Senior Deep Learning Research Engineer, LLM Inference
Core
Develop and optimize Large Language Model (LLM) inference algorithms to enhance NVIDIA's software, focusing on low-latency and high-throughput scenarios for global users.
Role type
Senior Deep Learning Research Engineer (LLM Inference)
Builds
NVIDIA software products and inference engines
Domain
AI/ML, High-Performance Computing, LLM Inference
Deliverable
production ML models
Required skills
Python, PyTorch, High-Performance Computing (HPC), GPU cluster management, algorithmic optimization, benchmarking, profiling, LLM architecture knowledge
Preferred skills
Publications in top-tier AI/ML conferences, experience with speculative decoding, parallelization strategies, vLLM, TensorRT-LLM
Technologies
PyTorch, vLLM, TensorRT-LLM, NVIDIA GPUs
Responsibilities
Develop and improve benchmarks and evaluation pipelines; Design experimental frameworks for algorithmic tradeoff analysis; Prototype new LLM inference algorithms; Profile performance on NVIDIA hardware to identify bottlenecks; Collaborate with cross-functional teams; Stay current with LLM inference research and translate advances into solutions
Seniority
Senior, hands-on IC