CareerPlanSign in

Research Engineer - Post-Training

USA or Australia🌐 Remote💼 Full-time🗓 2026-08-31 → 2026-09-26

Core

Build end-to-end RL post-training stacks for large language models running on decentralized consumer GPUs and Macs over the public internet.

Role type

Senior IC research engineer (RL post-training & distributed systems)

Builds

Decentralized RL training loops, reward computation pipelines, and policy update mechanisms for non-trusted, geo-distributed inference.

Domain

Decentralized AI, Protocol Learning, Large Language Models, Reinforcement Learning

Deliverable

production ML models

Required skills

RL post-training (RLHF, RLVR, reasoning RL), distributed systems engineering, asynchronous training loops, weight synchronization, Python, PyTorch

Preferred skills

Experience with slow networks or decentralized/federated setups, serving-engine internals (vLLM, SGLang), reward modeling, P2P networking, NAT traversal

Technologies

Python, PyTorch, vLLM, SGLang

Responsibilities

Design and implement the full RL training loop including rollout ingestion, reward computation, and policy updates; Adapt standard RL algorithms for high-latency, partially trusted, asynchronous environments; Build evaluation frameworks and release the first decentralized post-trained model artifacts.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.