Research Engineer - Distributed Training
Core
Building decentralized AI training orchestration and infrastructure for large-scale reinforcement learning and frontier models.
Role type
Senior Research Engineer (Distributed Training)
Builds
Decentralized training orchestration solutions, open-source libraries, and frameworks for distributed model training.
Domain
AI/ML Infrastructure, Distributed Systems, Reinforcement Learning
Deliverable
production ML models
Required skills
Distributed training techniques, compute & memory optimization, large-scale model training pipelines, MLOps, open-source framework development, technical writing
Preferred skills
Experience with frontier agentic models, async RL training, verifiable evals
Technologies
PyTorch Distributed, DeepSpeed, MosaicML's LLM Foundry, Ray
Responsibilities
Lead research for decentralized training orchestration, optimize performance and cost of AI workloads, contribute to open-source libraries, publish research in top-tier AI conferences, distill technical outcomes into blogs
Seniority
Senior, hands-on IC