高性能计算研发工程师-Ads Infra
Core
Optimize computation graph execution efficiency for ad/recommendation/search scenarios and build industry-leading high-performance training/inference engines.
Role type
Senior IC high-performance computing engineer (ML infrastructure)
Builds
High-performance training and inference engines, operator libraries, and distributed ML systems
Domain
AI infrastructure, GPU/NPU computing, distributed systems
Deliverable
production ML models
Required skills
C/C++/Python, Linux, data structures and algorithms, distributed system design, ML frameworks (PyTorch/TensorFlow/PaddlePaddle)
Preferred skills
CUDA/Triton, TensorRT/Cutlass, ML compilers (XLA/MLIR/TVM), distributed training frameworks (FSDP/DeepSpeed/Megatron)
Technologies
GPU, NPU, Linux, C++, Python, PyTorch, TensorFlow, PaddlePaddle, CUDA, Triton, TensorRT, Cutlass, XLA, MLIR, TVM, FSDP, DeepSpeed, Megatron
Responsibilities
Optimize model training/inference computation graphs; Develop high-performance operator libraries; Research and implement cutting-edge hardware and heterogeneous computing technologies; Design and maintain large-scale distributed systems.