Senior Deep Learning Software Engineer, Inference
Core
Design, build, and optimize GPU-accelerated software for high-performance open-source frameworks serving large-scale LLM and Generative AI models.
Role type
Senior IC deep learning inference software engineer
Builds
NVIDIA inference libraries (vLLM, SGLang, FlashInfer) and GPU-accelerated model serving pipelines
Domain
AI/ML infrastructure, Large Language Models, Generative AI, GPU computing
Deliverable
production ML models
Required skills
C/C++ programming, software design, performance optimization, GPU architecture knowledge, deep learning model inference
Preferred skills
CUDA programming, OAI Triton, CUTLASS, NCCL, NVSHMEM, Python, performance modeling and profiling
Technologies
CUDA, OAI Triton, CUTLASS, NCCL, NVSHMEM, vLLM, SGLang, FlashInfer, PyTorch
Responsibilities
Optimize DL models across NVIDIA accelerators (datacenter GPUs to edge SoCs), contribute code to open-source inference libraries, collaborate on cross-framework solutions
Seniority
Senior, hands-on IC
