AI加速软件资深研发工程师-芯片研发
Core
Lead software-hardware co-optimization for large language models on self-developed AI chips, focusing on distributed training/inference performance, operator optimization, and system architecture.
Role type
Senior IC AI software engineer (chip/hardware co-design)
Builds
High-throughput distributed training and inference systems for large models on custom AI hardware
Domain
AI hardware acceleration, distributed computing, high-performance computing
Deliverable
production ML models
Required skills
AI compiler development, operator optimization, high-performance communication (NCCL/RDMA), C/C++/Python, deep learning framework internals, system architecture design, performance modeling (roofline), cluster cost analysis
Preferred skills
LLM/multimodal model expertise, large-scale cluster training (thousands of cards), custom interconnect/NoC development, model quantization/sparse/distillation, AI server topology design
Responsibilities
Design and optimize distributed training/inference schemes for self-developed chips; implement and optimize high-performance computing and collective communication operators; explore and deploy model quantization and sparsity solutions; guide technical roadmap and mentor team members; drive cross-team collaboration for software stack evolution.
Seniority
Senior, hands-on IC with technical leadership
