Researcher, Alignment Interpretability
Core
Develop and publish research on techniques for understanding representations of deep networks and engineer infrastructure for studying model internals at scale to ensure AI alignment.
Role type
Researcher, Alignment Interpretability
Builds
Research publications and scalable infrastructure for mechanistic interpretability
Domain
Artificial Intelligence Safety, Deep Learning, Mechanistic Interpretability
Deliverable
production ML models | research
Required skills
mechanistic interpretability, AI safety & alignment, quantitative reasoning, research process, Python
Preferred skills
experience in AI safety & alignment, long-term AI safety thinking, curiosity about large-scale AI systems
Technologies
Python
Responsibilities
Develop and publish research on techniques for understanding representations of deep networks; Engineer infrastructure for studying model internals at scale; Collaborate across teams on unique OpenAI projects; Guide research directions toward demonstrable usefulness and/or long-term scalability
Seniority
Senior, hands-on IC