大语言模型可解释性研究员 - Seed Model
Core
Researching interpretability of large language models via large-scale sparse dictionary learning to understand internal neurons and circuits.
Role type
Research Scientist (LLM Interpretability)
Builds
Novel training and evaluation paradigms based on activation engineering for LLMs
Domain
Artificial Intelligence / Large Language Models / Machine Learning
Deliverable
research
Required skills
Large-scale sparse dictionary learning, neuron/circuit analysis, reasoning mechanisms, activation engineering, C/C++, Python, data structures, algorithms
Preferred skills
Published influential papers/projects in ML/LLMs/RL, deep research experience in multimodal or reinforcement learning
Technologies
C/C++, Python
Responsibilities
Analyze high-order neurons and circuits affecting model reasoning, personality, and role-playing; Explore activation-based training and evaluation methods
Seniority
Senior, hands-on IC