CareerPlanGet AI match score →

Staff Research Engineer, Post-training & Evaluation

🌐 Remote💼 Full-time💰 $230,000–$230,000🗓 2026-07-08 → 2026-07-31

Core

Define the science of model development feedback loops, establishing evaluation standards and post-training methodologies for Reddit-native Large Language Models.

Role type

Staff Research Engineer (LLM Post-Training & Evaluation)

Builds

Foundational LLMs powering Safety, Moderation, Search, and Ads

Domain

Internet / Large Language Models / AI Safety

Deliverable

production ML models

Required skills

evaluation reliability, statistical significance, custom evaluation harnesses, model-as-a-judge methodology, SFT recipe design, checkpoint selection, synthetic data generation, safety policy translation, loss curve diagnosis

Preferred skills

MLflow, fine-tuning frameworks (Axolotl, TorchTune), synthetic data techniques (Self-Instruct), preference optimization (DPO, RLHF, RLAIF, GRPO), multimodal model evaluation

Technologies

Python, Hugging Face Transformers, vLLM, lm-eval-harness, PyTorch, FSDP2, DeepSpeed ZeRO-3

Responsibilities

Define the "Reddit Benchmark" evaluation standard for model quality; Own evaluation reliability and statistical rigor; Design model-as-a-judge methodology; Set post-training recipes and strategy; Evaluate base and CPT checkpoints; Drive synthetic data generation strategy; Partner with Safety Engineering to translate policy into metrics; Diagnose post-training instability; Lead research direction and mentor engineers

Seniority

Staff, hands-on IC with strategic leadership

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗