CareerPlanSign in

Senior Deep Learning Software Engineer, Inference

6 Locations💼 Full-time💰 $152,000–$152,000🗓 2026-08-27 → 2026-09-25

Core

Design, build, and optimize GPU-accelerated software for high-performance open-source frameworks serving large-scale LLM and Generative AI models.

Role type

Senior IC deep learning inference software engineer

Builds

NVIDIA inference libraries (vLLM, SGLang, FlashInfer) and GPU-accelerated model serving pipelines

Domain

AI/ML infrastructure, Large Language Models, Generative AI, GPU computing

Deliverable

production ML models

Required skills

C/C++ programming, software design, performance optimization, GPU architecture knowledge, deep learning model inference

Preferred skills

CUDA programming, OAI Triton, CUTLASS, NCCL, NVSHMEM, Python, performance modeling and profiling

Technologies

CUDA, OAI Triton, CUTLASS, NCCL, NVSHMEM, vLLM, SGLang, FlashInfer, PyTorch

Responsibilities

Optimize DL models across NVIDIA accelerators (datacenter GPUs to edge SoCs), contribute code to open-source inference libraries, collaborate on cross-framework solutions

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.