CareerPlanGet AI match score →

Deep Learning Software Engineer, TensorRT Performance - New College Grad 2026

California, US💼 Full-time🗓 2026-04-04 → 2026-07-31

Core

Develop GPU-accelerated deep learning inference software like TensorRT to optimize and deploy models for Generative AI, Recommenders, and Vision.

Role type

Deep Learning Software Engineer (Inference Performance)

Builds

GPU-accelerated inference software (TensorRT, TensorRT-EdgeLLM, Torch-TensorRT) and optimized model pipelines.

Domain

Deep Learning Inference, GPU Acceleration, Generative AI

Deliverable

production ML models

Required skills

C++, Python, Deep Learning frameworks (PyTorch, JAX, TensorFlow, ONNX), Inference libraries (TensorRT, TensorRT-LLM, vLLM, SGLang, FlashInfer), Performance analysis, Performance optimization, Quantization, Scheduling, Memory management, Distributed inference, Graph compiler algorithms, Frontend operators, Code generators

Preferred skills

GPU architecture knowledge, OSS framework integration

Technologies

TensorRT, TensorRT-EdgeLLM, Torch-TensorRT, PyTorch, JAX, TensorFlow, ONNX, C++, Python

Responsibilities

Establish performance benchmarking methodologies and analysis workflows, Contribute features and code to NVIDIA/OSS inference frameworks, Develop new model pipelines with optimized performance, Scale performance of deep learning models across different NVIDIA accelerator architectures, Collaborate with cross-functional teams on generative AI, automotive, robotics, image understanding, and speech understanding

Seniority

Junior, New College Grad 2026

Sourced via adzuna · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Adzuna ↗