Research Scientist, Interpretability
Rewrite
## About Anthropic
Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.
## About the role:
When you see what modern language models are capable of, do you wonder, "How do these things work? How can we trust them?"
The Interpretability team at Anthropic is working to reverse-engineer how trained models work because we believe that a mechanistic understanding is the most robust way to make advanced systems safe. We’re looking for researchers and engineers to join our efforts.
People mean many different things by "interpretability". We're focused on mechanistic interpretability, which aims to discover how neural network parameters map to meaningful algorithms. Some useful analogies might be to think of us as trying to do "biology" or "neuroscience" of neural networks using "microscopes" we build, or as treating neural networks as binary computer programs we're trying to "reverse engineer".
A few places to learn more about our work and team at a high level are [this introduction to Interpretability](https://www.youtube.com/watch?v=TxhhMTOTMDg) from our research lead, [Chris Olah](https://colah.github.io/about.html); a [discussion of our work](https://open.spotify.com/episode/5UF79Uu94ia0fwC32a89LU) on the [Hard Fork podcast](https://www.nytimes.com/column/hard-fork) produced by the New York Times, and this [blog post](https://www.anthropic.com/research/engineering-challenges-interpretability) (and accompanying video) sharing more about some of the engineering challenges we’d had to solve to get these results. Some of our team's notable publications include [A Mathematical Framework for Transformer Circuits](https://transformer-circuits.pub/2021/framework/index.html), [In-context Learning and Induction Heads](https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html), [Toy Models of Superposition](https://transformer-circuits.pub/2022/toy_model/index.html), [Scaling Monosemanticity](https://transformer-circuits.pub/2024/scaling-monosemanticity/), and our Circuits' [Methods](https://transformer-circuits.pub/2025/attribution-graphs/methods.html) and [Biology](https://transformer-circuits.pub/2025/attribution-graphs/biology.html) papers. This work builds on ideas from members' work prior to Anthropic such as the [original circuits thread](https://distill.pub/2020/circuits/), [Multimodal Neurons](https://distill.pub/2021/multimodal-neurons/), [Activation Atlases](https://distill.pub/2019/activation-atlas/), and [Building Blocks](https://distill.pub/2018/building-blocks/).
We aim to create a solid foundation for mechanistically understanding neural networks and making them safe (see our [vision post](https://transformer-circuits.pub/2023/interpretability-dreams/index.html)). In the short term, we have focused on resolving the issue of "superposition" (see [Toy Models of Superposition](https://transformer-circuits.pub/2022/toy_model/index.html), [Superposition, Memorization, and Double Descent](https://transformer-circuits.pub/2023/toy-double-descent/index.html), and our [May 2023 update](https://transformer-circuits.pub/2023/may-update/index.html)), which causes the computational units of the models, like neurons and attention heads, to be individually uninterpretable,
## Responsibilities
- Reverse-engineer how trained models work to create a mechanistic understanding of neural networks.
- Investigate and resolve issues like "superposition" that hinder interpretability.
- Collaborate with researchers and engineers to build tools and frameworks for interpreting neural networks.
- Contribute to publications and share findings through blogs, videos, and podcasts.
- Work on foundational research to make advanced AI systems safe and beneficial for society.
## Requirements
- Strong background in machine learning, neural networks, and interpretability research.
- Experience with transformer models and related architectures.
- Proficiency in programming languages such as Python, with experience in deep learning frameworks like PyTorch or TensorFlow.
- Strong analytical and problem-solving skills.
- Excellent communication skills to collaborate effectively with cross-functional teams.
- Ability to work independently and manage multiple projects simultaneously.
## Nice to Have
- Experience with research in mechanistic interpretability or related fields.
- Familiarity with academic publishing and presenting research findings.
- Knowledge of computational neuroscience or related disciplines.
- Experience with large-scale AI systems and their challenges.
## Benefits
- Opportunity to work on cutting-edge AI research with a mission-driven team.
- Collaborative and innovative work environment.
- Access to resources and tools to support research and development.
- Potential for professional growth and impact on the future of AI.
Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.