CareerPlanSign in

Compute Architecture Software Engineer

China, Shanghai💼 Full-time🗓 2025-09-25 → 2026-09-25

Core

Develop and optimize software solutions to accelerate LLM inference using GPU technology for NVIDIA's TRTLLM project.

Role type

Senior IC LLM inference software engineer (GPU)

Builds

Optimized LLM inference software running on single PCs to multi-GPU clusters

Domain

AI / Deep Learning / GPU Computing

Deliverable

production ML models

Required skills

GPU programming, LLM inference, Python, C++, CUDA, deep learning frameworks

Preferred skills

None stated

Technologies

CUDA, Python, C++, deep learning frameworks

Responsibilities

Develop and optimize software to accelerate LLM inference using GPU technology; Collaborate with engineers to implement and refine GPU-based algorithms; Analyze methods to improve performance across diverse computing environments

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.