AI Inference Performance Engineer
Core
Optimizing and benchmarking GenAI inference performance on NVIDIA accelerators for language, video, and speech workloads.
Role type
Senior IC GPU performance engineer (inference)
Builds
High-performance inference tools and benchmarks for large-scale LLMs, VLMs, and diffusion models
Domain
GPU deep learning / High-performance computing
Deliverable
production ML models
Required skills
Python, C++, PyTorch, JAX, LLM/VLM architecture, CUDA programming, distributed inference, roofline analysis, quantization, scheduling, memory management
Preferred skills
TensorRT-LLM, vLLM, SGLang, kernel development (CUTLASS, Triton), compiler/runtime optimization, scale-out orchestration (MPI, NCCL, K8S), team leadership
Technologies
TensorRT-LLM, SGLang, vLLM, CUDA, PyTorch, JAX, NCCL, K8S, MPI
Responsibilities
Drive end-to-end optimization pipelines for inference; Define and optimize next-generation AI workloads; Architect distributed inference from single-GPU to rack-scale clusters; Establish performance methodology via profiling and roofline analysis; Contribute to open-source inference frameworks; Lead technical execution and team performance
Seniority
Senior, hands-on IC with leadership responsibilities