AI Infra工程师-Commercial AI
Core
Design and develop distributed training and inference infrastructure for LLM/VLM/AIGC models, optimizing compute, communication, and storage performance for large-scale clusters.
Role type
Senior IC AI Infrastructure Engineer (LLM/VLM/AIGC)
Builds
High-performance model training and inference engines, distributed cluster systems
Domain
Artificial Intelligence, Large Language Models, Distributed Systems
Deliverable
production ML models
Required skills
C++, Python, CUDA, Triton, PyTorch, TensorFlow, Data Structures & Algorithms, Concurrent Programming, Operator Fusion, Memory Management, Communication Overlap, Graph Compilation, Dynamic Batching, KV Cache Management, TensorRT-LLM, vLLM, LightX2V, FlashAttention, Quantization (FP8/INT8), Sparse Distillation
Preferred skills
Experience with NPU optimization, Deep learning performance acceleration techniques
Technologies
NVIDIA GPU, CUDA, Triton, NPU, PyTorch, TensorFlow, TensorRT-LLM, vLLM, LightX2V
Responsibilities
Optimize compute, communication, and storage performance for multi-thousand GPU clusters; Tune NVIDIA GPU and NPU performance via operator fusion and memory management; Design and implement low-latency, high-throughput inference engines; Research and implement latest industry infrastructure technologies.
Seniority
Senior, hands-on IC