Senior Deep Learning Software Engineer, TensorRT Performance
Core
Optimizing performance, deployment, and serving of deep learning inference solutions across NVIDIA's ecosystem including TensorRT and OSS frameworks.
Role type
Senior Deep Learning Software Engineer (Inference Performance)
Builds
GPU-accelerated deep learning inference software, model pipelines, and performance benchmarking workflows.
Domain
Deep Learning Inference, GPU Acceleration, High-Performance Computing
Deliverable
production ML models
Required skills
C++, Python, Deep Learning frameworks (PyTorch, JAX, TensorFlow, ONNX), Inference libraries (TensorRT, TensorRT-LLM, vLLM, SGLang, FlashInfer), Performance analysis, Performance optimization, Graph compiler algorithms, Frontend operators, Code generation, Quantization, Scheduling, Memory management, Distributed inference
Preferred skills
GPU architecture knowledge, Generative AI experience
Technologies
TensorRT, TensorRT-EdgeLLM, Torch-TensorRT, PyTorch, JAX, TensorFlow, ONNX, CUDA
Responsibilities
Establish performance benchmarking methodologies and analysis workflows; Contribute features and code to NVIDIA/OSS inference frameworks; Develop new model pipelines with optimized performance; Collaborate with cross-functional teams to develop innovative inference solutions; Scale performance of deep learning models across different NVIDIA accelerator architectures.
Seniority
Senior, hands-on IC