Senior AI Infrastructure Engineer - Model Training
Core
Design and optimize high-throughput data loading, streaming, and distributed training infrastructure for multimodal sensor data (camera, LiDAR, radar) to maximize GPU utilization for large-scale AI model training.
Role type
Senior IC AI Infrastructure Engineer (Model Training)
Builds
Distributed training clusters and scalable dataset construction pipelines for multimodal driving data
Domain
Autonomous driving / AI Infrastructure
Deliverable
production ML models
Required skills
Distributed training frameworks (PyTorch DDP/FSDP, DeepSpeed, Megatron), High-performance data pipelines (WebDataset, MosaicML Streaming), GPU optimization (mixed precision, kernel fusion, memory hierarchy), Systems programming (C++/CUDA/Triton), Profiling tools (Nsight, PyTorch Profiler)
Preferred skills
Experience with NVIDIA B200 accelerators, NVLink/InfiniBand interconnects
Technologies
PyTorch, NVIDIA B200, WebDataset, MosaicML Streaming, NCCL, NVLink, InfiniBand, C++, CUDA, Triton, BF16, FP8
Responsibilities
Design high-throughput data loading and streaming systems for multimodal sensor data; Build and optimize distributed training infrastructure across multi-node GPU clusters; Maximize utilization of modern accelerators through mixed-precision training and memory optimization; Profile end-to-end training pipelines to eliminate bottlenecks; Develop scalable dataset construction pipelines; Partner with ML teams to scale new architectures
Seniority
Senior, hands-on IC