CareerPlanSign in

百度公有云异构加速工程师(J91679)

北京市,上海市💼 Full-time🗓 2026-07-21 → 2026-09-28

Core

Develop and optimize training frameworks and inference acceleration for large language models and multimodal models on heterogeneous hardware.

Role type

Senior IC machine-learning engineer (LLM & multimodal inference)

Builds

Production-scale training clusters and optimized inference services for large language and multimodal models

Domain

Artificial Intelligence / Large Language Models / High-Performance Computing

Deliverable

production ML models

Required skills

PyTorch, CUDA, Megatron-LM, DeepSpeed, Verl, C++/Python mixed development, Ray, Transformer architecture, Diffusion models, MoE structures

Preferred skills

TPU, Ascend, AMD MI300 hardware porting, open-source framework contributions, thousand-card cluster optimization

Technologies

Megatron-LM, DeepSpeed, Verl, PD separation frameworks, PyTorch Profiler, Nsight Systems

Responsibilities

Optimize parallel strategies and memory usage in training frameworks; Develop and optimize inference acceleration strategies including quantization and speculative decoding; Develop operator adaptations and performance optimizations for new heterogeneous hardware; Analyze performance bottlenecks using profiling tools to deliver optimization solutions.

Sourced via baidu · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.