Staff Research Engineer, Post-training & Evaluation
Core
Define the science of model development feedback loops, establishing evaluation standards and post-training methodologies for Reddit-native Large Language Models.
Role type
Staff Research Engineer (LLM Post-Training & Evaluation)
Builds
Foundational LLMs powering Safety, Moderation, Search, and Ads
Domain
Internet / Large Language Models / AI Safety
Deliverable
production ML models
Required skills
evaluation reliability, statistical significance, custom evaluation harnesses, model-as-a-judge methodology, SFT recipe design, checkpoint selection, synthetic data generation, safety policy translation, loss curve diagnosis
Preferred skills
MLflow, fine-tuning frameworks (Axolotl, TorchTune), synthetic data techniques (Self-Instruct), preference optimization (DPO, RLHF, RLAIF, GRPO), multimodal model evaluation
Technologies
Python, Hugging Face Transformers, vLLM, lm-eval-harness, PyTorch, FSDP2, DeepSpeed ZeRO-3
Responsibilities
Define the "Reddit Benchmark" evaluation standard for model quality; Own evaluation reliability and statistical rigor; Design model-as-a-judge methodology; Set post-training recipes and strategy; Evaluate base and CPT checkpoints; Drive synthetic data generation strategy; Partner with Safety Engineering to translate policy into metrics; Diagnose post-training instability; Lead research direction and mentor engineers
Seniority
Staff, hands-on IC with strategic leadership