CareerPlanGet AI match score →

Software Engineer, cuDNN - Deep Learning

China, Shanghai💼 Full-time🗓 2025-09-18 → 2026-08-01

Core

Design, build, and ship cuDNN, a GPU-accelerated library of primitives for deep neural networks, including optimized large language model (LLM) support.

Role type

Senior IC software engineer (GPU-accelerated deep learning primitives)

Builds

NVIDIA's AI software stack (cuDNN library)

Domain

AI / Deep Learning / GPU Computing

Deliverable

production ML models

Required skills

C/C++ development, CUDA development, Python, linear algebra, high-level software architecture design, performance analysis, profiling, code optimization

Preferred skills

GPU programming and optimization expertise (CUDA/OpenCL), practical experience with machine learning (deep learning), computer architecture knowledge, MLIR development, compiler optimization

Technologies

CUDA, Python, C/C++, cuDNN, LLMs, MLIR

Responsibilities

Develop production-quality software for NVIDIA's AI software stack; Analyze performance of workloads and propose software improvements; Collaborate with deep learning engineers and GPU architects on applications like generative AI and autonomous driving; Contribute to API design, software architecture, performance modeling, testing, and GPU kernel development.

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗