机器学习生产管理研发工程师(编排调度/推理服务方向) - Data AML
Core
Manage the full lifecycle of online inference services for applied machine learning models, ensuring stability, cost optimization, and performance for recommendation, ad, CV, voice, and NLP systems.
Role type
Senior IC Machine Learning Infrastructure Engineer (Inference Serving & Orchestration)
Builds
Scalable, stable, and cost-efficient inference platforms for internal business units (Douyin, Toutiao) and external enterprise clients via Volcano Engine.
Domain
Internet / Applied Machine Learning / Cloud Infrastructure
Deliverable
production ML models
Required skills
Linux administration, Python, Go, Kubernetes (K8s) ecosystem, model inference optimization (quantization, pruning), GPU/TPU resource management, CI/CD automation, capacity planning, fault tolerance design.
Preferred skills
Experience with large-scale model serving architectures, performance tuning of heterogeneous hardware, serverless inference concepts, automated self-healing systems.
Responsibilities
Monitor and troubleshoot online inference services to ensure high availability; optimize GPU/CPU utilization and resource scheduling; implement elastic scaling based on traffic patterns; build automated CI/CD pipelines for service deployment; design and execute disaster recovery and multi-region failover strategies; tune underlying infrastructure and model parameters for maximum throughput.
Seniority
Senior, hands-on IC