AI加速软件资深研发工程师-芯片研发
Core
Evaluate and optimize large model performance on self-developed AI chips, leading hardware-software co-design for distributed training and inference.
Role type
Senior IC AI software engineer (chip optimization)
Builds
Distributed training/inference systems for large models (e.g., Doubao, Seedance) on custom hardware
Domain
AI hardware acceleration, high-performance computing, distributed systems
Deliverable
production ML models
Required skills
AI compiler development, operator optimization, high-performance communication, hardware-software co-design, C/C++/Python, deep learning framework internals, system architecture design
Preferred skills
LLM/multimodal model expertise, NCCL/DeepEP/RDMA knowledge, GPU/DSA/CUDA experience, model quantization/sparse/distillation tools, large-scale cluster deployment
Technologies
C, C++, Python, NCCL, DeepEP, RDMA, CUDA, NoC
Responsibilities
Design and optimize high-performance computing operators and collective communication primitives; explore and implement model quantization, sparsity, and distillation strategies; lead technical roadmap and mentor team members; drive cross-team collaboration for software stack evolution.
Seniority
Senior, hands-on IC with technical leadership
