ML Infra Engineer, Modeling
Core
Design, implement, and maintain large-scale ML training infrastructure to support foundation model development for physical robots and devices.
Role type
Senior IC ML Infrastructure Engineer (Modeling)
Builds
Scalable GPU/TPU training clusters, job orchestration systems, and reusable JAX training pipelines.
Domain
Robotics, Physical AI, Large-scale Foundation Models
Deliverable
production ML models
Required skills
Software engineering fundamentals, Large-scale training experience, JAX, Distributed training, Cloud platform management, Performance optimization, Experiment lifecycle management
Preferred skills
ML systems background, Hardware-level optimization, Robotics domain knowledge, Abstraction design
Technologies
JAX, PyTorch, TPU, GPU, Kubernetes, GCP, AWS, SLURM
Responsibilities
Design and maintain training/inference infrastructure systems, Scale distributed training across clusters, Optimize memory usage and device utilization, Build abstractions for experiment management, Partner with researchers to translate needs into infra capabilities, Contribute to core JAX model and training code
Seniority
Senior, hands-on IC