CareerPlanSign in

Helix AI Engineer, Training Performance

HQ💼 Full-time💰 $200,000–$200,000🗓 2026-08-14 → 2026-09-26

Core

Optimizing distributed training frameworks and GPU kernels for 100B+ parameter models across 100k+ GPUs to maximize hardware utilization for autonomous humanoid robots.

Role type

Senior IC AI Training Performance Engineer

Builds

High-performance distributed training infrastructure and custom GPU kernels for large-scale model training

Domain

Robotics, Large Language Models, High-Performance Computing

Deliverable

production ML models

Required skills

GPU architecture optimization, CUDA/C++ kernel development, distributed training frameworks (FSDP, context parallel), collective communication (NCCL, RDMA), performance profiling (Nsight, PyTorch Profiler), Python, fault tolerance engineering, hardware-efficiency metrics (MFU/HFU)

Preferred skills

Heterogeneous/multi-datacenter training orchestration, open-source ML contributions (PyTorch, DeepSpeed, JAX), non-NVIDIA accelerator experience (AMD, TPU)

Technologies

Triton, CUDA, PyTorch, NCCL, NVLink, InfiniBand, RDMA, FSDP, Gluon

Responsibilities

Optimize training performance for 100B+ parameter models across 100k+ GPUs; Write and optimize custom kernels (Triton/CUDA); Build tooling for continuous performance monitoring and regression detection; Optimize data loading and preprocessing pipelines; Improve checkpointing and fault tolerance; Partner with researchers to co-design model architectures; Extend kernel compilers; Evaluate emerging accelerator architectures; Explore model/data parallelisms

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.