CareerPlanSign in

大模型训练性能优化工程师(训练算子)(深圳/北京/上海/杭州)

Beijing, China💼 Full-time🗓 2026-09-28

Core

Design, implement, and optimize deep learning training operators for large-scale model training scenarios to maximize throughput, latency, and memory utilization.

Role type

Senior IC machine-learning engineer (training operator optimization)

Builds

Optimized training operators and communication schemes for 3D parallel training systems

Domain

AI Infrastructure / High-Performance Computing / GPU Systems

Deliverable

production ML models

Required skills

C/C++, CUDA programming, CUTLASS, Triton, GPU architecture (SM, Warp, Memory Hierarchy), 3D parallelism (Data/Tensor/Pipeline), performance profiling (Nsight, nvprof, perf)

Preferred skills

Experience with compiler libraries, network optimization, hardware architecture tracking

Technologies

CUDA, CUTLASS, Triton, Nsight, nvprof, perf

Responsibilities

Design and implement deep learning training operators; Perform end-to-end performance analysis and tuning for large model training; Design and optimize operators and communication schemes under 3D parallel training; Collaborate with distributed training and system teams to improve efficiency and stability; Track and integrate latest hardware and system software technologies.

Seniority

Senior, hands-on IC

Sourced via tencent · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.