大模型异构训练推理研发工程师(J96922)
Core
Engineering optimization of AIGC large language models for training and inference scenarios, ensuring stable operation on domestic computing platforms.
Role type
Senior IC machine-learning engineer (large model training & inference optimization)
Builds
Optimized large model training and inference pipelines on domestic GPU hardware
Domain
Artificial Intelligence / Large Language Models / Domestic Computing Hardware
Deliverable
production ML models
Required skills
Python, C/C++, PyTorch, Megatron-LM, vLLM, sglang, TensorRT-LLM, DeepSpeed, GPU architecture, distributed training (DP/TP/PP), memory management, operator optimization, low-precision training, LoRA
Preferred skills
Multi-card heterogeneous computing acceleration, KV Cache strategies, model compression (quantization/distillation), ZeRO/Offload, MoE model tuning, system-level debugging
Technologies
PyTorch, Megatron-LM, vLLM, sglang, TensorRT-LLM, DeepSpeed, domestic GPUs
Responsibilities
Optimize model training and inference performance (throughput, latency, resource utilization); adapt models to domestic GPU platforms; build evaluation and testing systems for performance and stability; lead customer project delivery and on-site support for deployment issues.
Seniority
Senior, hands-on IC