Researcher, Alignment CoT Monitorability
Core
Design and run empirical studies to measure and improve the monitorability of chain-of-thought in frontier reasoning models to support scalable AI oversight.
Role type
Researcher, Alignment CoT Monitorability
Builds
Evaluations for model monitorability, monitoring models/methods, and practical oversight recommendations for large training runs
Domain
AI Safety, Alignment, Large Language Models (LLMs), Model Training
Deliverable
production ML models | research
Required skills
empirical ML, training/evaluating/debugging large LLMs, experimental design, hypothesis formulation, data analysis, translating research to engineering
Preferred skills
chain-of-thought interpretability, alignment research, model behavior investigation, moving between research ideation and engineering execution
Technologies
LLMs, reinforcement learning, synthetic data, pre-training, mid-training, post-training
Responsibilities
Design and run empirical studies of chain-of-thought monitorability; Build evaluations to measure monitorability of high-stakes misbehavior; Investigate training interventions affecting monitorability; Analyze model behavior to generate hypotheses and recommendations; Translate findings into practical monitoring approaches; Collaborate across model training and alignment teams; Produce externally publishable research
Seniority
Individual Contributor, Researcher