Model Policy Manager, Agentic Safety
Core
Define behavioral policies, evaluations, and safeguards to ensure autonomous AI models behave safely in real-world environments over long horizons.
Role type
Senior IC AI safety policy engineer (agentic safety)
Builds
Behavioral policies, evaluation frameworks, monitoring systems, and safeguards for frontier AI models
Domain
AI safety, alignment, and autonomous agent security
Deliverable
production ML models | product features
Required skills
AI agent safety, adversarial mindset, threat modeling, empirical evaluation, data analysis, policy framework design, cross-functional collaboration
Preferred skills
cybersecurity, privacy, long-horizon model research
Technologies
evaluation datasets, training data, monitoring tools
Responsibilities
Identify vulnerabilities in model interactions with tools and external systems; develop threat models for misaligned behavior; build frameworks to understand harmful outcomes; translate findings into policy and evaluation criteria; develop human data campaigns and gold sets; partner with research and engineering teams; inform deployment decisions and safety reports; build post-deployment monitoring approaches