Deep Learning Software Engineer, TensorRT Performance - New College Grad 2026
Core
Develop GPU-accelerated deep learning inference software like TensorRT to optimize and deploy models for Generative AI, Recommenders, and Vision.
Role type
Deep Learning Software Engineer (Inference Performance)
Builds
GPU-accelerated inference software (TensorRT, TensorRT-EdgeLLM, Torch-TensorRT) and optimized model pipelines.
Domain
Deep Learning Inference, GPU Acceleration, Generative AI
Deliverable
production ML models
Required skills
C++, Python, Deep Learning frameworks (PyTorch, JAX, TensorFlow, ONNX), Inference libraries (TensorRT, TensorRT-LLM, vLLM, SGLang, FlashInfer), Performance analysis, Performance optimization, Quantization, Scheduling, Memory management, Distributed inference, Graph compiler algorithms, Frontend operators, Code generators
Preferred skills
GPU architecture knowledge, OSS framework integration
Technologies
TensorRT, TensorRT-EdgeLLM, Torch-TensorRT, PyTorch, JAX, TensorFlow, ONNX, C++, Python
Responsibilities
Establish performance benchmarking methodologies and analysis workflows, Contribute features and code to NVIDIA/OSS inference frameworks, Develop new model pipelines with optimized performance, Scale performance of deep learning models across different NVIDIA accelerator architectures, Collaborate with cross-functional teams on generative AI, automotive, robotics, image understanding, and speech understanding
Seniority
Junior, New College Grad 2026