大模型训推框架技术架构师(J102502)
Core
Design and evolve training and inference frameworks for large language models, optimizing performance for MoE and reinforcement learning architectures while building AI infrastructure.
Role type
Senior IC machine-learning infrastructure architect
Builds
Self-developed and open-sourced training frameworks for multi-modal and frontier models; large-scale distributed inference systems; domestic heterogeneous AI chip clusters
Domain
AI infrastructure, large language models, domestic semiconductor ecosystem
Deliverable
production ML models
Required skills
Large-scale distributed training optimization, high-performance inference engine development, operator optimization, heterogeneous computing, MoE architecture support, multi-modal model training, domestic AI chip adaptation, software ecosystem construction
Preferred skills
Team management and talent development, cross-functional coordination between chip, framework, and model teams
Technologies
MoE, reinforcement learning, multi-modal models, domestic AI chips, distributed systems, KV storage systems
Responsibilities
Lead architecture design and technical planning for LLM training/inference frameworks; optimize inference engines and cluster scheduling to reduce latency and costs; enable and adapt frameworks for domestic AI chip clusters to support frontier models