Researcher, Alignment Science
Core
Design and run experiments to train models to understand user intent, act faithfully, verify behavior, and honestly report limitations using scalable alignment methods.
Role type
Research Engineer / Research Scientist (Alignment Science)
Builds
Scalable methods for instruction-following, honesty, and robustness in frontier AI models
Domain
Artificial Intelligence / Machine Learning / Alignment Science
Deliverable
production ML models
Required skills
hands-on training and debugging of large LLMs, Python, PyTorch, reinforcement learning, post-training, preference optimization, scalable oversight, model evaluation, mathematical rigor, quantitative experimentation
Preferred skills
competitive programming, math contests, systems work, technical problem solving
Technologies
Python, PyTorch, reinforcement learning frameworks
Responsibilities
Design and implement alignment experiments for intent following, honesty, calibration, and robustness; Train and evaluate models using reinforcement learning and empirical ML methods; Develop evaluations for failure modes like hallucination and reward hacking; Study methods for model self-verification and honest reporting; Build monitoring and inference-time interventions; Investigate scaling of alignment methods with model capability and compute; Integrate successful techniques into training and deployment workflows; Produce externally publishable research
Seniority
Senior, hands-on IC