Research Engineer - Post-Training
Core
Build end-to-end RL post-training stacks for large language models running on decentralized consumer GPUs and Macs over the public internet.
Role type
Senior IC research engineer (RL post-training & distributed systems)
Builds
Decentralized RL training loops, reward computation pipelines, and policy update mechanisms for non-trusted, geo-distributed inference.
Domain
Decentralized AI, Protocol Learning, Large Language Models, Reinforcement Learning
Deliverable
production ML models
Required skills
RL post-training (RLHF, RLVR, reasoning RL), distributed systems engineering, asynchronous training loops, weight synchronization, Python, PyTorch
Preferred skills
Experience with slow networks or decentralized/federated setups, serving-engine internals (vLLM, SGLang), reward modeling, P2P networking, NAT traversal
Technologies
Python, PyTorch, vLLM, SGLang
Responsibilities
Design and implement the full RL training loop including rollout ingestion, reward computation, and policy updates; Adapt standard RL algorithms for high-latency, partially trusted, asynchronous environments; Build evaluation frameworks and release the first decentralized post-trained model artifacts.
Seniority
Senior, hands-on IC