Research Engineer, Large-Scale Training
Core
Build and optimize large-scale training infrastructure for open foundation models to enable efficient fine-tuning for downstream applications.
Role type
Senior IC research engineer (large-scale ML training systems)
Builds
Production training infrastructure, experimental infrastructure, and open-source model support
Domain
Artificial Intelligence / Machine Learning Systems / Distributed Training
Deliverable
production ML models | infrastructure
Required skills
Python, PyTorch, multi-GPU/multi-node training, distributed training paradigms (data/tensor/pipeline/expert parallelism), GPU architecture, mixed-precision training, performance profiling, bottleneck elimination, experiment design, CUDA/Triton, NCCL/NVSHMEM, FSDP/DeepSpeed/Megatron-LM
Preferred skills
Optimized GPU kernel development, large-scale experiment management, open-source contributions, ML product operations
Technologies
PyTorch, CUDA, Triton, NCCL, NVSHMEM, FSDP, DeepSpeed, Megatron-LM
Responsibilities
Design and optimize core components of large-scale training infrastructure; Integrate new model architectures and validate training correctness; Profile distributed workloads to eliminate bottlenecks; Execute experiments to benchmark new approaches; Partner with scientists to productionize novel training methods; Enable support for newly released open-source foundation models; Build and maintain experimental infrastructure for research and production
Seniority
Senior, hands-on IC