Researcher, Misalignment Research
Core
Designing and executing cutting-edge attacks, building adversarial evaluations, and advancing understanding of how AI safety measures fail to identify and quantify future AGI misalignment risks.
Role type
Senior Researcher, AI Safety & Misalignment
Builds
Automated infrastructure for red-teaming and stress testing; rigorous, repeatable evaluations of dangerous capabilities.
Domain
Artificial Intelligence / Machine Learning Safety
Deliverable
production ML models | research | infrastructure
Required skills
AI red-teaming, adversarial ML, security research, system-level stress testing, automated tool development, failure mode analysis, cross-functional collaboration, technical writing/publishing
Preferred skills
Experience with large-scale codebases, modern ML techniques, mentoring
Technologies
Large-scale codebases, evaluation infrastructure, automated testing frameworks
Responsibilities
Design and implement worst-case demonstrations of AGI alignment risks; Develop adversarial and system-level evaluations; Create automated tools to scale red-teaming; Conduct research on failure modes of alignment techniques; Publish influential papers; Partner with engineering, policy, and legal teams; Mentor engineers and researchers
Seniority
Senior, hands-on IC with mentorship responsibilities