Research Infrastructure Engineer, Training Systems
Core
Building systems layer infrastructure to turn novel ML research ideas into runnable, measurable training workloads for large models.
Role type
Senior IC research infrastructure engineer (training systems)
Builds
Infrastructure for large-scale model training and experimentation
Domain
AI research / distributed systems / ML training
Deliverable
production ML models
Required skills
distributed systems, Python, PyTorch, GPU programming, networking, storage systems, API design, performance optimization, debugging from traces/logs
Preferred skills
systems instincts, clean abstractions, empathy for researcher workflows, evidence-based debugging
Technologies
Python, PyTorch, GPUs
Responsibilities
Build and maintain infrastructure for large-scale model training and experimentation; Design APIs and interfaces for complex training workflows; Improve reliability, debuggability, and performance across training and data pipelines; Debug issues spanning Python, PyTorch, distributed systems, GPUs, networking, and storage; Write tests, benchmarks, and diagnostics that catch meaningful regressions
Seniority
Senior, hands-on IC