Senior Deep Learning Software Engineer, LLM Performance
Core
Optimizing LLM inference performance, deployment, and serving across NVIDIA accelerators using GPU-accelerated software frameworks.
Role type
Senior Deep Learning Software Engineer (LLM Performance)
Builds
High-performance LLM inference solutions, TensorRT LLM, VLLM, SGLang, and Triton frameworks.
Domain
Deep Learning, Generative AI, GPU Computing, High-Performance Computing
Deliverable
production ML models
Required skills
Python, C, C++, PyTorch, JAX, TensorFlow, CUDA, performance modeling, profiling, kernel development, software design
Preferred skills
LLM framework experience, DL compiler experience, performance optimization of HPC applications, CPU/GPU architecture knowledge
Technologies
TensorRT, VLLM, SGLang, Triton, CUDA, PyTorch, JAX, TensorFlow
Responsibilities
Optimize LLM/VLM/GenAI models for inference and serving; Scale performance across NVIDIA accelerator architectures; Contribute code to NVIDIA/OSS LLM frameworks; Collaborate on generative AI and multimodal solutions.
Seniority
Senior, hands-on IC