Safeguards Enforcement Analyst, Cyber Harm
Core
Review flagged content and accounts to detect and mitigate misuse of AI systems for malicious cyber operations, malware development, and offensive exploitation.
Role type
Safeguards Enforcement Analyst (Cyber Harm)
Builds
Safe AI products and services by enforcing policies against cyber threats
Domain
AI Safety / Cybersecurity
Deliverable
production ML models | product features
Required skills
cybersecurity knowledge (offensive techniques, exploit development, malware analysis, vulnerability research), content review and abuse investigation, SQL/Python for data analysis, stakeholder communication, generative AI prompt engineering
Preferred skills
trust & safety/abuse investigation/cybersecurity investigations/threat intelligence in tech/AI, large language model misuse understanding, abuse monitoring program experience, product policy implementation at scale, government agency/regulated environment experience
Technologies
SQL, Python, generative AI products
Responsibilities
Review flagged content and accounts to make accurate, well-documented enforcement decisions; Detect and mitigate potential misuse of AI systems to facilitate cyberattacks, malware creation, and exploitation tooling; Triage and escalate novel, ambiguous, or high-severity cases; Provide detailed feedback to policy design teams on gaps; Partner with Engineering and Data Science teams to improve detection model precision and recall; Maintain high accuracy and consistency standards across review queues
Seniority
Mid-level IC