CareerPlanSign in

Software Engineer, AI and DL Kernel Libraries

China, Shanghai💼 Full-time🗓 2026-06-24 → 2026-09-26

Core

Design, build, and optimize low-level GPU kernels and inference runtimes for NVIDIA's AI software stack, serving high-performance workloads for LLMs, generative AI, and autonomous driving.

Role type

Senior IC systems software engineer (AI inference kernels & runtimes)

Builds

Production-quality AI software stack including cuDNN, FlashInfer, and LLM inference runtimes

Domain

AI Systems / GPU Computing / Deep Learning Infrastructure

Deliverable

production ML models

Required skills

C/C++, Python, CUDA, deep learning frameworks (PyTorch, JAX, TensorFlow, ONNX), linear algebra, performance profiling, software abstraction design

Preferred skills

GPU kernel development, JIT compilation, code generation, MLIR/Apache TVM/TensorIR, GPU performance modeling, open-source contributions

Technologies

CUDA, cuDNN, FlashInfer, vLLM, SGLang, TensorRT-LLM, Triton, cuTile, MLIR, Apache TVM, TensorIR

Responsibilities

Develop production-quality software for NVIDIA's AI stack; Design and optimize kernels for LLM inference and generative AI; Build JIT compilation and code generation systems; Analyze workload performance and tune software; Collaborate with GPU architecture and compiler teams; Contribute to open-source inference ecosystems

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.