Helix AI Engineer, Training Performance
Core
Optimizing distributed training frameworks and GPU kernels for 100B+ parameter models across 100k+ GPUs to maximize hardware utilization for autonomous humanoid robots.
Role type
Senior IC AI Training Performance Engineer
Builds
High-performance distributed training infrastructure and custom GPU kernels for large-scale model training
Domain
Robotics, Large Language Models, High-Performance Computing
Deliverable
production ML models
Required skills
GPU architecture optimization, CUDA/C++ kernel development, distributed training frameworks (FSDP, context parallel), collective communication (NCCL, RDMA), performance profiling (Nsight, PyTorch Profiler), Python, fault tolerance engineering, hardware-efficiency metrics (MFU/HFU)
Preferred skills
Heterogeneous/multi-datacenter training orchestration, open-source ML contributions (PyTorch, DeepSpeed, JAX), non-NVIDIA accelerator experience (AMD, TPU)
Technologies
Triton, CUDA, PyTorch, NCCL, NVLink, InfiniBand, RDMA, FSDP, Gluon
Responsibilities
Optimize training performance for 100B+ parameter models across 100k+ GPUs; Write and optimize custom kernels (Triton/CUDA); Build tooling for continuous performance monitoring and regression detection; Optimize data loading and preprocessing pipelines; Improve checkpointing and fault tolerance; Partner with researchers to co-design model architectures; Extend kernel compilers; Evaluate emerging accelerator architectures; Explore model/data parallelisms
Seniority
Senior, hands-on IC