CareerPlanGet AI match score →

Research Scientist, Safety Post Training

San Francisco, CA💼 Full-time💰 $216,000–$216,000🗓 2026-07-09 → 2026-07-31

Core

Develop and apply post-training methods and interpretability techniques to make frontier AI systems safer and better understood by researchers and policymakers.

Role type

Research Scientist (AI Safety & Post-Training)

Builds

Post-training pipelines, interpretability-informed evaluations, and safety standards/benchmarks for frontier AI models.

Domain

Artificial Intelligence / Machine Learning Safety

Deliverable

production ML models

Required skills

Post-training methods, RLHF, DPO, GRPO, Machine Learning research, Generative AI, Model safety evaluation, Interpretability techniques, Cross-functional collaboration

Preferred skills

Mechanistic interpretability, Probing, Red-teaming, Adversarial evaluation, Reward hacking analysis, Sycophancy detection, Alignment faking detection

Technologies

RLHF, DPO, GRPO

Responsibilities

Design and run post-training pipelines to study training choices' effect on safety and alignment; Develop interpretability-informed evaluations to reveal unsafe behaviors; Collaborate with policymakers and engineers to translate findings into safety standards and benchmarks.

Seniority

Mid-Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗