北京-自动驾驶大规模AI系统优化与异构计算工程师(J100724)
Core
Develop inference operators and optimize distributed training frameworks for large-scale autonomous driving models on NVIDIA GPUs.
Role type
Senior IC AI infrastructure engineer (high-performance computing)
Builds
High-throughput distributed training clusters and optimized inference engines for autonomous driving
Domain
Autonomous driving + High-performance computing + Deep learning infrastructure
Deliverable
production ML models
Required skills
C/C++, CUDA, GPU architecture, distributed systems, kernel fusion, NCCL, Triton, vLLM, Megatron, TVM, TensorRT
Preferred skills
Deep learning framework core development, open-source community contributions (PyTorch, Triton)
Technologies
NVIDIA GPU, Triton, vLLM, Megatron, NCCL, TVM, TensorRT, PyTorch
Responsibilities
Develop and optimize inference operators for large models on NVIDIA GPUs; Design and optimize distributed training frameworks for thousand-node clusters; Optimize memory usage and communication performance for large model structures; Track and deploy latest AI infrastructure technologies
Seniority
Senior, hands-on IC