AI Infra 研发工程师(J104808)
Core
Build industry-leading AI heterogeneous computing container platforms and optimize large model training/inference for AIGC, data centers, finance, and energy clients.
Role type
Senior IC AI Infrastructure Engineer (Heterogeneous Computing & Model Optimization)
Builds
High-performance, high-stability container platforms for AI application deployment and optimized training/inference solutions.
Domain
AI Infrastructure, Heterogeneous Computing, Large Language Models (LLMs)
Deliverable
production ML models
Required skills
Golang, Python, C++, CUDA, Kubernetes, PyTorch, Distributed Training, Model Optimization, Container Orchestration
Preferred skills
GPU Architecture, OpenCL, FlashAttention, MoE Architecture, Zero/Offload, SOTA Frameworks (Megatron, DeepSpeed, vLLM)
Technologies
Kubernetes, PyTorch, CUDA, OpenCL, Megatron, DeepSpeed, vLLM, SGLang
Responsibilities
Develop and tune operators for self-developed chips to improve model performance and accuracy; Explore and apply innovations in task management, resource scheduling, and distributed training for large-scale heterogeneous clusters; Optimize training/inference efficiency using SOTA principles; Adapt common large models to custom hardware; Contribute to open-source machine learning frameworks.
Seniority
Senior, hands-on IC