Research Program Manager - Research Infrastructure
Core
Scale research infrastructure to support massive frontier-scale AI model training runs across pre-training, mid-training, and post-training phases.
Role type
Senior Research Program Manager (Infrastructure)
Builds
Reliable, high-performance training environments and cluster reliability for frontier AI models
Domain
Artificial Intelligence / Large-scale Distributed Systems
Deliverable
production ML models
Required skills
Cross-functional program ownership, incident triage and escalation management, stakeholder management with technical ICs and leadership, process creation for cross-team handoffs, technical fluency in distributed training frameworks and GPU clusters, crisis management
Preferred skills
Experience in ML/AI or large-scale distributed systems environments, ability to operate in high-ambiguity settings, experience building processes from zero to one
Technologies
Megatron, GPU clusters, distributed training frameworks, schedulers, networking, storage systems
Responsibilities
Own cross-functional programs spanning training infrastructure and cluster reliability; Drive end-to-end coordination scaling the training stack; Jump into active incidents to triage and drive resolution; Partner with engineering leads to identify bottlenecks and define priorities; Build visibility into training run health and infrastructure performance; Create lightweight processes for config management and checkpoint workflows
Seniority
Senior, hands-on IC