CareerPlanGet AI match score →

Training Performance Engineer

San Francisco💼 Full-time🗓 2025-10-16 → 2026-07-31

Core

Drive efficiency improvements across a distributed machine-learning training runtime by analyzing large-scale training runs, identifying utilization gaps, and designing optimizations for throughput and uptime.

Role type

Senior IC systems performance engineer (distributed ML training)

Builds

High-performance, fault-tolerant training frameworks and distributed process management for frontier-scale model runs

Domain

Artificial Intelligence / Distributed Systems / High-Performance Computing

Deliverable

production ML models

Required skills

GPU kernel performance analysis, collective communication optimization, I/O bottleneck investigation, model sharding, distributed system debugging, Python, C++, PyTorch, JAX, TensorFlow

Preferred skills

Rust, CUDA, NCCL, MPI, UCX, large-scale data loading, checkpointing systems, ML compiler optimization

Technologies

PyTorch, JAX, TensorFlow, NCCL, MPI, UCX

Responsibilities

Profile end-to-end training runs to identify performance bottlenecks across compute, communication, and storage; Optimize GPU utilization and throughput for large-scale distributed model training; Collaborate with runtime and systems engineers to improve kernel efficiency, scheduling, and collective communication performance; Implement model graph transforms to improve end-to-end throughput; Build tooling to monitor and visualize MFU, throughput, and uptime across clusters; Partner with researchers to ensure new model architectures scale efficiently during pre-training; Contribute to infrastructure decisions that improve reliability and efficiency of large training jobs

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗