Research Scientist, Safety Post Training
Core
Develop and apply post-training methods and interpretability techniques to make frontier AI systems safer and better understood by researchers and policymakers.
Role type
Research Scientist (AI Safety & Post-Training)
Builds
Post-training pipelines, interpretability-informed evaluations, and safety standards/benchmarks for frontier AI models.
Domain
Artificial Intelligence / Machine Learning Safety
Deliverable
production ML models
Required skills
Post-training methods, RLHF, DPO, GRPO, Machine Learning research, Generative AI, Model safety evaluation, Interpretability techniques, Cross-functional collaboration
Preferred skills
Mechanistic interpretability, Probing, Red-teaming, Adversarial evaluation, Reward hacking analysis, Sycophancy detection, Alignment faking detection
Technologies
RLHF, DPO, GRPO
Responsibilities
Design and run post-training pipelines to study training choices' effect on safety and alignment; Develop interpretability-informed evaluations to reveal unsafe behaviors; Collaborate with policymakers and engineers to translate findings into safety standards and benchmarks.
Seniority
Mid-Senior, hands-on IC