AI计算基础设施工程师 - 基础技术
Core
Design, build, and operate GPU-centric compute, network, and storage systems for large-scale model training and inference.
Role type
Senior IC AI infrastructure engineer
Builds
Kubernetes-based compute scheduling and resource management platforms
Domain
Cloud infrastructure, AI for Infra, GPU computing
Deliverable
production ML models
Required skills
Kubernetes cluster management, Linux kernel development, GPU driver development, performance tuning, system virtualization, model parallelism strategies
Preferred skills
LLM architecture knowledge, DPU development, Agent Infra experience
Technologies
Kubernetes, Linux, GPU accelerators
Responsibilities
Plan and optimize AI compute infrastructure for training and inference; enhance system performance and resource utilization via hardware topology optimization; build multi-tenant resource management platforms; perform cost and reliability optimization for AI workloads.