CareerPlanSign in

Research Engineer - Pre-training

USA or Australia🌐 Remote💼 Full-time🗓 2026-08-31 → 2026-09-26

Core

Build the distributed training system for Protocol Learning, enabling large model pre-training on heterogeneous consumer-grade devices connected via the internet.

Role type

Senior Research Engineer (Distributed Systems & ML)

Builds

Scalable, fault-tolerant distributed training infrastructure for frontier-scale models

Domain

Decentralized AI, Protocol Learning, Distributed Machine Learning

Deliverable

production ML models

Required skills

Distributed training (PyTorch, FSDP, DeepSpeed, Megatron), Model parallelism (data, pipeline, tensor), Production-quality Python, Concurrency, Failure handling, Profiling

Preferred skills

Large language model training (Nemotron, Qwen, OLMo), P2P networking, NAT traversal, Post-training and RL, Inference and serving systems

Technologies

PyTorch, FSDP, DeepSpeed, Megatron

Responsibilities

Implement and optimize model-parallel training across heterogeneous GPUs under low-bandwidth, high-latency links; Reduce communication overhead while maintaining model convergence; Ensure run survival through node churn via robust checkpointing and state synchronization; Build monitoring systems for throughput, bottlenecks, and model quality across hundreds of devices

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.