CareerPlanGet AI match score →

Research Infrastructure Engineer, Training Systems

San Francisco💼 Full-time🗓 2026-04-27 → 2026-07-31

Core

Building systems layer infrastructure to turn novel ML research ideas into runnable, measurable training workloads for large models.

Role type

Senior IC research infrastructure engineer (training systems)

Builds

Infrastructure for large-scale model training and experimentation

Domain

AI research / distributed systems / ML training

Deliverable

production ML models

Required skills

distributed systems, Python, PyTorch, GPU programming, networking, storage systems, API design, performance optimization, debugging from traces/logs

Preferred skills

systems instincts, clean abstractions, empathy for researcher workflows, evidence-based debugging

Technologies

Python, PyTorch, GPUs

Responsibilities

Build and maintain infrastructure for large-scale model training and experimentation; Design APIs and interfaces for complex training workflows; Improve reliability, debuggability, and performance across training and data pipelines; Debug issues spanning Python, PyTorch, distributed systems, GPUs, networking, and storage; Write tests, benchmarks, and diagnostics that catch meaningful regressions

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗