Research Engineer - Pre-training
Core
Build the distributed training system for Protocol Learning, enabling large model pre-training on heterogeneous consumer-grade devices connected via the internet.
Role type
Senior Research Engineer (Distributed Systems & ML)
Builds
Scalable, fault-tolerant distributed training infrastructure for frontier-scale models
Domain
Decentralized AI, Protocol Learning, Distributed Machine Learning
Deliverable
production ML models
Required skills
Distributed training (PyTorch, FSDP, DeepSpeed, Megatron), Model parallelism (data, pipeline, tensor), Production-quality Python, Concurrency, Failure handling, Profiling
Preferred skills
Large language model training (Nemotron, Qwen, OLMo), P2P networking, NAT traversal, Post-training and RL, Inference and serving systems
Technologies
PyTorch, FSDP, DeepSpeed, Megatron
Responsibilities
Implement and optimize model-parallel training across heterogeneous GPUs under low-bandwidth, high-latency links; Reduce communication overhead while maintaining model convergence; Ensure run survival through node churn via robust checkpointing and state synchronization; Build monitoring systems for throughput, bottlenecks, and model quality across hundreds of devices
Seniority
Senior, hands-on IC