后端调度编排工程师 - Data AML
Core
Designing and optimizing distributed scheduling frameworks and training/inference runtimes for large-scale recommendation, advertising, CV, and NLP models.
Role type
Senior IC backend scheduling and orchestration engineer (ML infrastructure)
Builds
Distributed training runtimes, inference architectures, and scheduling frameworks for recommendation and advertising systems.
Domain
Machine Learning Infrastructure / Distributed Systems
Deliverable
production ML models
Required skills
Go, Python, Linux, Kubernetes, distributed system principles, resource modeling, AutoScaling, multi-cloud scheduling
Preferred skills
Experience with vLLM, Ray, TFX, RL, Finetuning, model distillation
Technologies
Kubernetes, Godel, Yarn, Flink, MapReduce, Mesos, Celery, veRL, vLLM, Ray, TFX
Responsibilities
Optimize cluster utilization and scheduling strategies for distributed frameworks; design elastic distributed training runtimes for large-scale models; build robust online inference architectures for recommendation systems.
Seniority
Senior, hands-on IC