CareerPlanGet AI match score →

Software Engineer, AI and DL Kernel Libraries

China, Shanghai💼 Full-time🗓 2026-06-24 → 2026-07-30

Core

Design, build, and optimize low-level GPU kernels and inference runtimes for NVIDIA's AI software stack, serving high-performance workloads for LLMs, generative AI, and autonomous driving.

Role type

Senior IC systems software engineer (AI inference kernels & runtimes)

Builds

Production-quality AI software stack including cuDNN, FlashInfer, and LLM inference runtimes

Domain

AI Systems / GPU Computing / Deep Learning Infrastructure

Deliverable

production ML models

Required skills

C/C++, Python, CUDA, deep learning frameworks (PyTorch, JAX, TensorFlow, ONNX), linear algebra, performance profiling, software abstraction design

Preferred skills

GPU kernel development, JIT compilation, code generation, MLIR/Apache TVM/TensorIR, GPU performance modeling, open-source contributions

Technologies

CUDA, cuDNN, FlashInfer, vLLM, SGLang, TensorRT-LLM, Triton, cuTile, MLIR, Apache TVM, TensorIR

Responsibilities

Develop production-quality software for NVIDIA's AI stack; Design and optimize kernels for LLM inference and generative AI; Build JIT compilation and code generation systems; Analyze workload performance and tune software; Collaborate with GPU architecture and compiler teams; Contribute to open-source inference ecosystems

Seniority

Senior, hands-on IC

Rewrite
## Responsibilities - Develop production-quality software that ships as part of NVIDIA's AI software stack, including cuDNN, FlashInfer, and optimized support for large language model inference workloads. - Innovate and develop new AI systems technologies for efficient inference, with a focus on performance, scalability, maintainability, and usability. - Design, implement, and optimize kernels for high-impact AI workloads across LLM inference, generative AI, computer vision, autonomous driving, and recommender systems. - Design and implement extensible software abstractions for deep learning libraries, LLM serving engines, and runtime systems. - Build and improve just-in-time compilation, code generation, and runtime technologies for performance-critical GPU workloads. - Analyze workload performance, tune current software, and propose improvements to future software and hardware-software interfaces. - Collaborate closely with engineers across deep learning frameworks, libraries, kernels, compilers, and GPU architecture teams at NVIDIA. - Contribute to open-source communities and ecosystem integrations where relevant, including projects such as FlashInfer, vLLM, and SGLang. ## Requirements - Master's degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience. - 3+ years of relevant industry, research, or systems software development experience in machine learning, deep learning systems, compilers, or GPU software. More experience is expected for senior-level candidates. - Strong programming skills in C/C++ and Python, with hands-on experience developing high-performance software. - Solid experience with CUDA development and GPU programming fundamentals. - Strong experience developing or using deep learning frameworks such as PyTorch, JAX, TensorFlow, or ONNX. - Good understanding of linear algebra, performance analysis, profiling, and code optimization. - Experience designing software abstractions, APIs, or higher-level system architecture for performance-sensitive systems. - Familiarity with modern machine learning and inference system trends, especially around LLMs and generative AI. - For senior candidates, strong experience in GPU kernel development and performance optimization, especially using CUDA C/C++, cuTile, Triton, or similar technologies, is expected. ## Nice to Have - Hands-on experience with inference engines and runtimes such as vLLM, SGLang, MLC, TensorRT-LLM, or similar systems. - Background in domain-specific compiler, code generation, or library solutions for LLM inference and training. - Expertise in machine learning compilers or IR systems such as MLIR, Apache TVM, TensorIR, or related technologies. - Practical experience with GPU performance modeling, computer architecture, or accelerator-oriented software design. - Open-source project ownership or meaningful contributions in deep learning systems, compilers, kernels, or inference infrastructure. ## Benefits - Opportunity to work on cutting-edge AI technologies and contribute to NVIDIA's AI software stack. - Collaboration with world-class engineers across deep learning software, compilers, GPU architecture, and open-source inference ecosystems. - Impact on real-world workloads at scale through performance-critical software development.
Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗