CareerPlanSign in

ML Systems Engineer, Large-Scale Model Training & RL Infrastructure

Palo Alto💼 Full-time💰 $195,200–$195,200🗓 2026-07-22 → 2026-09-27

Core

Building and maintaining distributed training infrastructure for large-scale SFT, pretraining, preference optimization, and RL workloads to enable reliable, reproducible, and efficient frontier model improvement.

Role type

Senior ML Systems Engineer (Large-Scale Model Training & RL Infrastructure)

Builds

Distributed training infrastructure, RL pipelines, rollout/reward model serving systems, and experiment orchestration tools.

Domain

AI Infrastructure / Distributed Systems / GPU Clusters

Deliverable

production ML models | infrastructure

Required skills

Python, PyTorch, distributed model training, GPU cluster workloads, transformer training bottlenecks, debugging production training jobs, quantitative reasoning for throughput/utilization/memory

Preferred skills

Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, Slurm, Kubernetes, RL infrastructure frameworks (verl, slime, AReaL, OpenRLHF), NCCL, CUDA, Triton, Nsight, InfiniBand, RDMA, RoCE, H100/H200/B200 clusters

Technologies

Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, verl, slime, AReaL, OpenRLHF, NCCL, CUDA, Triton, Nsight, InfiniBand, RDMA, RoCE, H100, H200, B200, Slurm, Kubernetes

Responsibilities

Build and maintain distributed training infrastructure for SFT, continued pretraining, preference optimization, and RL workloads; Integrate and extend frameworks such as Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, verl, slime, AReaL, OpenRLHF; Implement and debug parallelism strategies including tensor, pipeline, sequence/context, expert, and data parallelism; Build reliable rollout, reward model serving, replay/data buffer, checkpointing, evaluation, and experiment orchestration components for RL training; Profile and improve GPU utilization, communication efficiency, memory usage, and training throughput; Diagnose failures across NCCL, CUDA, PyTorch, Ray, schedulers, storage, networking, and checkpointing layers; Create reproducible training runs, launch scripts, dashboards, runbooks, and operational tooling for research users; Partner with research scientists to turn algorithmic training recipes into scalable, debuggable systems; Write clear design docs, incident reports, benchmark reports, and operating guides.

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.