公有云模型训推加速工程师(J104555)
Core
Optimizing performance for large model training (SFT & RL) and inference acceleration.
Role type
Senior IC machine-learning engineer (LLM training & inference optimization)
Builds
Optimized training pipelines and high-performance inference systems for large language models
Domain
Cloud computing + Large Language Models
Deliverable
production ML models
Required skills
CUDA programming, distributed training strategies, model quantization, operator fusion, performance profiling, PyTorch internals
Preferred skills
DeepSpeed, Megatron-LM, vLLM, TensorRT-LLM, SGLang, Nsight Systems, Cutlass DSL
Technologies
CUDA, PyTorch, NCCL, FlashAttention, GEMM, INT8/INT4/FP8, BF16
Responsibilities
Design and implement distributed training strategies (data, tensor, pipeline parallelism); Optimize inference via quantization, operator fusion, KV Cache, and speculative decoding; Profile and tune GPU hardware for speed and memory efficiency; Evaluate and integrate best practices from industry frameworks
Seniority
Senior, hands-on IC