百度公有云异构加速工程师(J91679)
Core
Develop and optimize training frameworks and inference acceleration for large language models and multimodal models on heterogeneous hardware.
Role type
Senior IC machine-learning engineer (LLM & multimodal inference)
Builds
Production-scale training clusters and optimized inference services for large language and multimodal models
Domain
Artificial Intelligence / Large Language Models / High-Performance Computing
Deliverable
production ML models
Required skills
PyTorch, CUDA, Megatron-LM, DeepSpeed, Verl, C++/Python mixed development, Ray, Transformer architecture, Diffusion models, MoE structures
Preferred skills
TPU, Ascend, AMD MI300 hardware porting, open-source framework contributions, thousand-card cluster optimization
Technologies
Megatron-LM, DeepSpeed, Verl, PD separation frameworks, PyTorch Profiler, Nsight Systems
Responsibilities
Optimize parallel strategies and memory usage in training frameworks; Develop and optimize inference acceleration strategies including quantization and speculative decoding; Develop operator adaptations and performance optimizations for new heterogeneous hardware; Analyze performance bottlenecks using profiling tools to deliver optimization solutions.
