CareerPlanSign in

Developer Technology Engineer - AI

China, Shanghai💼 Full-time🗓 2026-07-15 → 2026-09-25

Core

Build and optimize core parallel algorithms and data structures for GPUs to solve performance bottlenecks in large language model (LLM) training and inference, contributing to frameworks like Megatron, TRTLLM, SGLang, and vLLM.

Role type

Senior IC GPU software engineer (LLM optimization)

Builds

Production GPU kernels, communication libraries, and open-source frameworks for distributed LLM workloads

Domain

Artificial Intelligence / High-Performance Computing / GPU Architecture

Deliverable

production ML models

Required skills

C/C++/Python/Fortran, parallel programming, GPU programming, algorithm design, distributed systems, compiler optimization, linear algebra, numerical methods

Preferred skills

LLM framework optimization, distributed communication optimization, NVLink/InfiniBand/RoCE protocols, open-source contribution

Technologies

Megatron, TRTLLM, SGLang, vLLM, cuDNN, cuBLAS, CUTLASS, DeepGEMM, FlashMLA, FlashAttention, Flashinfer, NCCL, NVSHMEM, DeepEP

Responsibilities

Optimize GPU kernels and instruction-level tuning for LLM workloads; Refine distributed training and inference communication libraries; Collaborate with architecture and research teams to influence next-generation software platforms; Engage in deep optimization of high-performance operators and compiler settings

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.