Software Engineer, AI and DL Kernel Libraries
Core
Design, build, and optimize low-level GPU kernels and inference runtimes for NVIDIA's AI software stack, serving high-performance workloads for LLMs, generative AI, and autonomous driving.
Role type
Senior IC systems software engineer (AI inference kernels & runtimes)
Builds
Production-quality AI software stack including cuDNN, FlashInfer, and LLM inference runtimes
Domain
AI Systems / GPU Computing / Deep Learning Infrastructure
Deliverable
production ML models
Required skills
C/C++, Python, CUDA, deep learning frameworks (PyTorch, JAX, TensorFlow, ONNX), linear algebra, performance profiling, software abstraction design
Preferred skills
GPU kernel development, JIT compilation, code generation, MLIR/Apache TVM/TensorIR, GPU performance modeling, open-source contributions
Technologies
CUDA, cuDNN, FlashInfer, vLLM, SGLang, TensorRT-LLM, Triton, cuTile, MLIR, Apache TVM, TensorIR
Responsibilities
Develop production-quality software for NVIDIA's AI stack; Design and optimize kernels for LLM inference and generative AI; Build JIT compilation and code generation systems; Analyze workload performance and tune software; Collaborate with GPU architecture and compiler teams; Contribute to open-source inference ecosystems
Seniority
Senior, hands-on IC