深圳-AI异构计算工程师(J101237)
Core
Build and optimize distributed training frameworks and inference engines for large-scale foundation and specialized models (text, image, voice) to enable business deployment.
Role type
Senior IC AI Heterogeneous Computing Engineer (LLM Infrastructure)
Builds
High-performance distributed training frameworks, optimized inference engines, and RL toolchains for trillion-parameter models.
Domain
Artificial Intelligence / Large Language Models / Heterogeneous Computing
Deliverable
production ML models
Required skills
C/C++/Python, CUDA programming, operator optimization, distributed training frameworks (Megatron, DeepSpeed), inference frameworks (vLLM, SGlang), model compression, knowledge distillation, SFT, RL, GPU cluster performance analysis.
Preferred skills
Experience with trillion-parameter model deployment, speculative decoding, KV cache optimization, Agentic workflows.
Responsibilities
Optimize SFT, RL, model compression, and inference acceleration for business deployment; build and optimize distributed training frameworks for Post-train and distillation; optimize inference engines for TTFT and throughput; design RL toolchains for efficient iteration.
Seniority
Senior, hands-on IC