Abuse Investigator (AI Self-Improvement Risk)
Core
Investigate model behaviors exhibiting autonomous or agentic patterns (chaining, persistence, tool use) to identify safety risks and improve safeguards.
Role type
Senior IC abuse investigator (AI self-improvement risk)
Builds
Detection signals and tracking strategies for emerging agentic risk patterns
Domain
AI safety, security, and trust & safety
Deliverable
production ML models | dashboards & analysis
Required skills
technical investigations, SQL, Python, multi-step system analysis, threat analysis, failure mode identification, automated detection development
Preferred skills
experience in AI safety/security/cyber/trust & safety, presenting analytic work in technical/policy settings
Technologies
SQL, Python
Responsibilities
Review leads and investigate model behavior for agentic patterns, detect and analyze multi-step planning and capability chaining, develop signals to identify emerging risks, identify gaps in safeguards and propose improvements, communicate findings to stakeholders
Seniority
Senior, hands-on IC