Training / AI Infrastructure
Core
Optimize foundation model training stack performance by profiling bottlenecks, designing distributed training systems, and implementing low-level code for multi-node GPU clusters.
Role type
Senior IC ML infrastructure engineer (training systems)
Builds
Scalable, robust distributed training systems for foundation models
Domain
AI/ML infrastructure, High-Performance Computing (HPC)
Deliverable
production ML models
Required skills
Distributed systems, Python, CUDA/cuDNN/Triton, GPU kernel optimization, multi-node cluster management, memory management, data throughput optimization, networking optimization, monitoring/debugging tools development
Preferred skills
Experience with PyTorch, model parallelism, pipeline parallelism, context parallelism, hardware-software interaction tuning
Technologies
PyTorch, CUDA, cuDNN, Triton
Responsibilities
Profile and eliminate bottlenecks across data pipelines to GPU kernels, design and optimize distributed training systems for multi-node GPU clusters, implement efficient low-level code and integrate into high-level frameworks, optimize workloads for hardware efficiency, develop monitoring and debugging tools for large-scale runs
Seniority
Senior, hands-on IC