Research Scientist, Interpretability
Core
Reverse-engineer how trained neural network models work to build a mechanistic understanding for safety and alignment.
Role type
Research Scientist, Mechanistic Interpretability
Builds
Production language models (Sonnet, Haiku 3.5) and interpretability tools
Domain
Artificial Intelligence / Large Language Models / Mechanistic Interpretability
Deliverable
research
Required skills
scientific research, experimental design, Python programming, code writing, result interpretation
Preferred skills
prior work in interpretability, team science, clear communication
Technologies
Python
Responsibilities
Develop methods for understanding LLMs by reverse engineering algorithms learned in their weights; Design and run robust experiments in toy scenarios and at scale; Create and analyze new interpretability features and circuits; Build infrastructure for running experiments and visualizing results; Communicate results internally and publicly
Seniority
Individual Contributor, Research Scientist