CareerPlanSign in

Senior Deep Learning Software Engineer, Inference

2 Locations💼 Full-time💰 $184,000–$184,000🗓 2026-09-10 → 2026-09-25

Core

Design, build, and optimize GPU-accelerated software for high-performance deep learning frameworks (SGLang, vLLM) to serve large-scale LLM and Generative AI models.

Role type

Senior IC deep learning inference software engineer

Builds

High-performance inference libraries and model serving pipelines for NVIDIA accelerators

Domain

Artificial Intelligence / High-Performance Computing / GPU Software

Deliverable

production ML models

Required skills

C/C++ programming, software design, deep learning model optimization, GPU architecture knowledge, multi-GPU communication, CUDA programming, performance profiling

Preferred skills

Python, training DL models in production, building products for enterprise customers, experience with PyTorch, NCCL, NVSHMEM, OAI Triton, CUTLASS, FlashInfer

Technologies

CUDA, NCCL, NVSHMEM, OAI Triton, CUTLASS, SGLang, vLLM, FlashInfer, PyTorch, NVIDIA GPUs

Responsibilities

Optimize DL models across LLM, Multimodal, and Generative AI domains; Scale performance across different NVIDIA accelerator architectures; Contribute features and code to inference libraries; Collaborate on cross-framework optimization solutions

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.