Senior AI Infrastructure Engineer - Model Training
Core
Designing and optimizing high-throughput infrastructure to train large-scale AI models for autonomous trucking using massive multimodal sensor datasets.
Role type
Senior AI Infrastructure Engineer (Model Training)
Builds
Distributed training infrastructure and data pipelines for autonomous vehicle models
Domain
Autonomous ground transportation / AI Systems
Deliverable
production ML models
Required skills
Distributed training frameworks (PyTorch DDP/FSDP, DeepSpeed, Megatron), High-performance data pipelines (WebDataset, MosaicML), GPU optimization (mixed precision, kernel fusion), Systems programming (C++/CUDA/Triton), Profiling tools (Nsight, PyTorch Profiler)
Preferred skills
Experience with NVIDIA B200 accelerators, NVLink/InfiniBand interconnects
Technologies
PyTorch, DeepSpeed, Megatron, NCCL, WebDataset, MosaicML Streaming, NVIDIA B200, NVLink, InfiniBand, C++, CUDA, Triton, Python
Responsibilities
Design high-throughput data loading and streaming systems for multimodal sensor data; Build and optimize distributed training infrastructure across multi-node GPU clusters; Maximize utilization of modern accelerators through mixed-precision training and memory optimization; Profile end-to-end training pipelines to eliminate bottlenecks; Develop scalable dataset construction pipelines; Partner with ML teams to scale new architectures
Seniority
Senior, hands-on IC