CareerPlanGet AI match score →

Senior Deep Learning Software Engineer, TensorRT Performance

California, US💼 Full-time🗓 2026-03-24 → 2026-07-31

Core

Optimizing performance, deployment, and serving of deep learning inference solutions across NVIDIA's ecosystem including TensorRT and OSS frameworks.

Role type

Senior Deep Learning Software Engineer (Inference Performance)

Builds

GPU-accelerated deep learning inference software, model pipelines, and performance benchmarking workflows.

Domain

Deep Learning Inference, GPU Acceleration, High-Performance Computing

Deliverable

production ML models

Required skills

C++, Python, Deep Learning frameworks (PyTorch, JAX, TensorFlow, ONNX), Inference libraries (TensorRT, TensorRT-LLM, vLLM, SGLang, FlashInfer), Performance analysis, Performance optimization, Graph compiler algorithms, Frontend operators, Code generation, Quantization, Scheduling, Memory management, Distributed inference

Preferred skills

GPU architecture knowledge, Generative AI experience

Technologies

TensorRT, TensorRT-EdgeLLM, Torch-TensorRT, PyTorch, JAX, TensorFlow, ONNX, CUDA

Responsibilities

Establish performance benchmarking methodologies and analysis workflows; Contribute features and code to NVIDIA/OSS inference frameworks; Develop new model pipelines with optimized performance; Collaborate with cross-functional teams to develop innovative inference solutions; Scale performance of deep learning models across different NVIDIA accelerator architectures.

Seniority

Senior, hands-on IC

Sourced via adzuna · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Adzuna ↗