CareerPlanGet AI match score →

Safeguards Enforcement Analyst, Cyber Harm

San Francisco, CA💼 Full-time💰 $285,000–$285,000🗓 2026-07-10 → 2026-07-31

Core

Review flagged content and accounts to detect and mitigate misuse of AI systems for malicious cyber operations, malware development, and offensive exploitation.

Role type

Safeguards Enforcement Analyst (Cyber Harm)

Builds

Safe AI products and services by enforcing policies against cyber threats

Domain

AI Safety / Cybersecurity

Deliverable

production ML models | product features

Required skills

cybersecurity knowledge (offensive techniques, exploit development, malware analysis, vulnerability research), content review and abuse investigation, SQL/Python for data analysis, stakeholder communication, generative AI prompt engineering

Preferred skills

trust & safety/abuse investigation/cybersecurity investigations/threat intelligence in tech/AI, large language model misuse understanding, abuse monitoring program experience, product policy implementation at scale, government agency/regulated environment experience

Technologies

SQL, Python, generative AI products

Responsibilities

Review flagged content and accounts to make accurate, well-documented enforcement decisions; Detect and mitigate potential misuse of AI systems to facilitate cyberattacks, malware creation, and exploitation tooling; Triage and escalate novel, ambiguous, or high-severity cases; Provide detailed feedback to policy design teams on gaps; Partner with Engineering and Data Science teams to improve detection model precision and recall; Maintain high accuracy and consistency standards across review queues

Seniority

Mid-level IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗