CareerPlanSign in

ML Research Engineer, Training

San Francisco, CA💼 Full-time🗓 2026-09-20 → 2026-09-25

Core

Build end-to-end training stacks and large-scale data systems to turn terabyte-scale real-world robot data into capable models for home and business robots.

Role type

Senior ML Research Engineer (Training & Infrastructure)

Builds

Distributed training pipelines, data ingestion/transformation/storage systems, and experiment orchestration for robot learning models.

Domain

Robotics, Machine Learning, Cloud Infrastructure

Deliverable

production ML models

Required skills

PyTorch, JAX, distributed training (FSDP/DDP), GPU performance engineering, CUDA profiling, data pipeline optimization, loss curve analysis, experiment reproducibility

Preferred skills

Robot learning (VLAs, world models, RL), Kubernetes, SFT/RL fine-tuning, TensorRT, CUDA kernel development, video dataset handling

Technologies

PyTorch, JAX, CUDA, Nsight Systems, NCCL, Kubernetes, GCP, AWS, TensorRT

Responsibilities

Build distributed training stacks (data loading, checkpointing, orchestration); Develop high-throughput data systems for multimodal robot data; Optimize GPU utilization and data pipeline bottlenecks; Drive sampling and curation strategies; Profile and fix training instability and data bugs; Transition research prototypes to production infrastructure

Seniority

Senior, hands-on IC

Sourced via dover · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.