大模型AI Infra平台研发工程师(J105494)
Core
Design and develop large-scale AI infrastructure platforms unifying pre-training, mid-training, SFT, and RL scenarios to enable full observability, tracking, and unified views for metrics, logs, and resources.
Role type
Senior IC AI Infrastructure Platform Engineer
Builds
Large-scale AI training platforms and Agent-friendly APIs/workflows
Domain
AI Infrastructure / Cloud Native / Distributed Systems
Deliverable
production ML models
Required skills
Go, Python, C++, Kubernetes, Ray, distributed systems, MLOps/LLMOps, data versioning, experiment tracking, model deployment, agent tool calling, long-running task management
Preferred skills
System architecture design, cross-team coordination, complex problem decomposition
Technologies
Kubernetes, Ray, Go, Python, C++, Java
Responsibilities
Design and develop large-scale AI infrastructure platforms; build heterogeneous resource scheduling systems with fault tolerance and elasticity; construct Agent-friendly APIs and workflows; independently solve complex system problems and drive cross-team consensus
Seniority
Senior, hands-on IC