上海-AI Infra工程师(J100732)
Core
Design and optimize infrastructure for large model training and inference, including distributed frameworks, inference engines, and model compression/deployment.
Role type
Senior IC AI Infrastructure Engineer
Builds
High-throughput, stable training and inference systems for large models on heterogeneous hardware (GPU/NPU)
Domain
AI Infrastructure / High-Performance Computing
Deliverable
production ML models
Required skills
Distributed systems, Parallel computing, CUDA programming, C++, Python, Deep learning frameworks (PyTorch/PaddlePaddle), Model quantization, Inference optimization, Memory management
Preferred skills
Heterogeneous hardware optimization, Operator fusion, CUDA/XLA optimization
Responsibilities
Design and optimize distributed training frameworks and inference engines; Implement mixed precision, quantization (INT8/FP16), and KV Cache optimizations; Collaborate with algorithm teams to ensure efficient model execution on GPU/NPU; Architect AI systems to improve throughput and stability; Apply cutting-edge AI Infra techniques like operator fusion to production systems