CareerPlanSign in

Senior Machine Learning Engineer - Orchestration

USA💼 Full-time🗓 2026-09-14 → 2026-09-25

Core

Design and optimize distributed scheduling, autoscaling, and training runtimes for ultra-large-scale ML workloads including embeddings, recommendations, and reinforcement learning.

Role type

Senior IC machine learning engineer (orchestration & distributed systems)

Builds

Distributed training runtimes, computing APIs, online inference architectures, and ML operations platforms

Domain

Cloud infrastructure, distributed systems, machine learning

Deliverable

production ML models | infrastructure

Required skills

Go, Python, Linux, distributed scheduling frameworks (Kubernetes, Godel, YARN, Mesos, Celery), large-scale distributed systems design

Preferred skills

PyTorch, TensorFlow, AI infrastructure, hardware/software co-design, HPC, ML hardware architecture (GPUs, accelerators), training orchestration systems (veRL, vLLM, Ray, TFX)

Technologies

Kubernetes, Godel, YARN, Mesos, Celery, PyTorch, TensorFlow, veRL, vLLM, Ray, TFX

Responsibilities

Optimize distributed scheduling and orchestration strategies; Develop and integrate autoscaling and resource preemption; Build distributed training runtimes for ultra-large-scale embeddings; Design distributed computing APIs for recommendation and advertising workflows; Improve diagnosability of distributed training platforms; Build distributed online model inference architecture

Seniority

Senior, hands-on IC

Sourced via codingjobboard · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.