ML Research Engineer, Training
Core
Build end-to-end training stacks and large-scale data systems to turn terabyte-scale real-world robot data into capable models for home and business robots.
Role type
Senior ML Research Engineer (Training & Infrastructure)
Builds
Distributed training pipelines, data ingestion/transformation/storage systems, and experiment orchestration for robot learning models.
Domain
Robotics, Machine Learning, Cloud Infrastructure
Deliverable
production ML models
Required skills
PyTorch, JAX, distributed training (FSDP/DDP), GPU performance engineering, CUDA profiling, data pipeline optimization, loss curve analysis, experiment reproducibility
Preferred skills
Robot learning (VLAs, world models, RL), Kubernetes, SFT/RL fine-tuning, TensorRT, CUDA kernel development, video dataset handling
Technologies
PyTorch, JAX, CUDA, Nsight Systems, NCCL, Kubernetes, GCP, AWS, TensorRT
Responsibilities
Build distributed training stacks (data loading, checkpointing, orchestration); Develop high-throughput data systems for multimodal robot data; Optimize GPU utilization and data pipeline bottlenecks; Drive sampling and curation strategies; Profile and fix training instability and data bugs; Transition research prototypes to production infrastructure
Seniority
Senior, hands-on IC