CareerPlanGet AI match score →

Researcher, Alignment CoT Monitorability

San Francisco💼 Full-time🗓 2026-06-29 → 2026-07-31

Core

Design and run empirical studies to measure and improve the monitorability of chain-of-thought in frontier reasoning models to support scalable AI oversight.

Role type

Researcher, Alignment CoT Monitorability

Builds

Evaluations for model monitorability, monitoring models/methods, and practical oversight recommendations for large training runs

Domain

AI Safety, Alignment, Large Language Models (LLMs), Model Training

Deliverable

production ML models | research

Required skills

empirical ML, training/evaluating/debugging large LLMs, experimental design, hypothesis formulation, data analysis, translating research to engineering

Preferred skills

chain-of-thought interpretability, alignment research, model behavior investigation, moving between research ideation and engineering execution

Technologies

LLMs, reinforcement learning, synthetic data, pre-training, mid-training, post-training

Responsibilities

Design and run empirical studies of chain-of-thought monitorability; Build evaluations to measure monitorability of high-stakes misbehavior; Investigate training interventions affecting monitorability; Analyze model behavior to generate hypotheses and recommendations; Translate findings into practical monitoring approaches; Collaborate across model training and alignment teams; Produce externally publishable research

Seniority

Individual Contributor, Researcher

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗