Senior Deep Learning Software Engineer, Inference
Core
Design, build, and optimize GPU-accelerated software for high-performance deep learning frameworks (SGLang, vLLM) to serve large-scale LLM and Generative AI models.
Role type
Senior IC deep learning inference software engineer
Builds
High-performance inference libraries and model serving pipelines for NVIDIA accelerators
Domain
Artificial Intelligence / High-Performance Computing / GPU Software
Deliverable
production ML models
Required skills
C/C++ programming, software design, deep learning model optimization, GPU architecture knowledge, multi-GPU communication, CUDA programming, performance profiling
Preferred skills
Python, training DL models in production, building products for enterprise customers, experience with PyTorch, NCCL, NVSHMEM, OAI Triton, CUTLASS, FlashInfer
Technologies
CUDA, NCCL, NVSHMEM, OAI Triton, CUTLASS, SGLang, vLLM, FlashInfer, PyTorch, NVIDIA GPUs
Responsibilities
Optimize DL models across LLM, Multimodal, and Generative AI domains; Scale performance across different NVIDIA accelerator architectures; Contribute features and code to inference libraries; Collaborate on cross-framework optimization solutions
Seniority
Senior, hands-on IC
