MLOps技术专家 - Seed Model
Core
Ensuring stability and performance optimization for large-scale generative AI training and inference systems (LLM, T2I, T2V) within a global AI research team.
Role type
Senior MLOps Engineer (Generative AI Infrastructure)
Builds
Stable training and inference systems for large language models and generative media applications.
Domain
Generative AI, Large Scale Distributed Systems, Cloud Infrastructure
Deliverable
production ML models
Required skills
Python, Go, Linux, Shell scripting, Kubernetes, RDMA networking, High-performance storage, Resource scheduling, Generative AI system stability, Capacity planning
Preferred skills
LLM engineering, Multi-cloud management, Distributed training frameworks, Network performance tuning, Cost optimization
Responsibilities
Optimize training/inference system stability and performance; Design resource scheduling systems; Plan and maintain high-speed networks and heterogeneous compute clusters; Manage capacity delivery and compute efficiency; Build monitoring and fault recovery mechanisms.
Seniority
Senior, hands-on IC