CareerPlanSign in

Research Engineer, Large-Scale Training

San Francisco💼 Full-time💰 $200,000–$200,000🗓 2026-07-30 → 2026-09-26

Core

Build and optimize large-scale training infrastructure for open foundation models to enable efficient fine-tuning for downstream applications.

Role type

Senior IC research engineer (large-scale ML training systems)

Builds

Production training infrastructure, experimental infrastructure, and open-source model support

Domain

Artificial Intelligence / Machine Learning Systems / Distributed Training

Deliverable

production ML models | infrastructure

Required skills

Python, PyTorch, multi-GPU/multi-node training, distributed training paradigms (data/tensor/pipeline/expert parallelism), GPU architecture, mixed-precision training, performance profiling, bottleneck elimination, experiment design, CUDA/Triton, NCCL/NVSHMEM, FSDP/DeepSpeed/Megatron-LM

Preferred skills

Optimized GPU kernel development, large-scale experiment management, open-source contributions, ML product operations

Technologies

PyTorch, CUDA, Triton, NCCL, NVSHMEM, FSDP, DeepSpeed, Megatron-LM

Responsibilities

Design and optimize core components of large-scale training infrastructure; Integrate new model architectures and validate training correctness; Profile distributed workloads to eliminate bottlenecks; Execute experiments to benchmark new approaches; Partner with scientists to productionize novel training methods; Enable support for newly released open-source foundation models; Build and maintain experimental infrastructure for research and production

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.