CareerPlanSign in

Software Engineer, LLM Inference

China, Shanghai💼 Full-time🗓 2026-09-11 → 2026-09-25

Core

Develop robust, scalable inferencing software for LLMs and deep learning models, optimizing performance across multiple platforms.

Role type

Senior IC CPU computing engineer (LLM inference)

Builds

GPU-accelerated libraries (CUDA, cuDNN, TensorRT) and TensorRT Edge LLM features

Domain

AI-City and self-driving car solutions using GPU-accelerated computing

Deliverable

production ML models

Required skills

C/C++ programming, software design, performance analysis, debugging, test design, deep learning frameworks (PyTorch), LLM/generative model awareness

Preferred skills

Academic awareness of AI developments, proactive problem solving

Technologies

CUDA, cuDNN, TensorRT, TensorRT Edge, PyTorch

Responsibilities

Craft and develop robust inferencing software scaled to multiple platforms, perform performance analysis and optimization, update TensorRT and TensorRT Edge LLM based on academic developments, collaborate with software, research, and product teams

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.