Senior Software Engineer, Machine Learning Inference
Core
Designing and implementing inference software optimizations to power AI applications on NVIDIA GPUs, specifically for TensorRT and TensorRT-LLM.
Role type
Senior Software Engineer (Machine Learning Inference)
Builds
Deep learning inference software for datacenter, workstations, and PCs
Domain
AI / Machine Learning / GPU Computing
Deliverable
production ML models
Required skills
C++, CUDA, Python, Rust, Deep Learning Frameworks, Compilers, System Software
Preferred skills
Inference backends, GPU programming, LLM inference frameworks, close-to-metal performance analysis
Technologies
TensorRT, TensorRT-LLM, vLLM, SGLang, PyTorch, JAX, OpenCL
Responsibilities
Design and optimize TensorRT and TensorRT-LLM for inference applications; Develop C++, Python, and CUDA software for LLM and Generative AI deployment; Collaborate with deep learning experts and GPU architects on hardware and software design.
Seniority
Senior, hands-on IC