CareerPlanGet AI match score →

Senior Deep Learning Software Engineer, Inference

California, US💼 Full-time💰 $184,000–$184,000🗓 2026-05-07 → 2026-07-31

Core

Design, build, and optimize GPU-accelerated software for high-performance deep learning frameworks (SGLang, vLLM) to serve large-scale language and generative AI models.

Role type

Senior Deep Learning Software Engineer (Inference)

Builds

NVIDIA inference libraries (vLLM, SGLang, FlashInfer) and model serving pipelines for LLMs and Generative AI.

Domain

Artificial Intelligence / High-Performance Computing / GPU Accelerators

Deliverable

production ML models

Required skills

C/C++ programming, software design, deep learning inference optimization, multi-GPU communications, CUDA programming, performance profiling and debugging, CPU/GPU architecture knowledge

Preferred skills

Python, training DL models, performance modeling, building products for enterprise customers, open-source contributions (PyTorch)

Technologies

CUDA, NCCL, NVSHMEM, OAI Triton, CUTLASS, SGLang, vLLM, FlashInfer, PyTorch

Responsibilities

Performance optimization and tuning of DL models across LLM, Multimodal, and Generative AI domains; Scaling DL model performance across NVIDIA accelerator architectures; Contributing features and code to NVIDIA inference libraries; Collaborating with cross-functional teams on inference optimization solutions

Seniority

Senior, hands-on IC

Sourced via adzuna · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Adzuna ↗