大模型训练性能优化工程师(训练算子)(深圳/北京/上海/杭州)
Core
Design, implement, and optimize deep learning training operators for large-scale model training scenarios to maximize throughput, latency, and memory utilization.
Role type
Senior IC machine-learning engineer (training operator optimization)
Builds
Optimized training operators and communication schemes for 3D parallel training systems
Domain
AI Infrastructure / High-Performance Computing / GPU Systems
Deliverable
production ML models
Required skills
C/C++, CUDA programming, CUTLASS, Triton, GPU architecture (SM, Warp, Memory Hierarchy), 3D parallelism (Data/Tensor/Pipeline), performance profiling (Nsight, nvprof, perf)
Preferred skills
Experience with compiler libraries, network optimization, hardware architecture tracking
Technologies
CUDA, CUTLASS, Triton, Nsight, nvprof, perf
Responsibilities
Design and implement deep learning training operators; Perform end-to-end performance analysis and tuning for large model training; Design and optimize operators and communication schemes under 3D parallel training; Collaborate with distributed training and system teams to improve efficiency and stability; Track and integrate latest hardware and system software technologies.
Seniority
Senior, hands-on IC