CareerPlanSign in

Safeguards Enforcement Analyst, User Well-being

San Francisco, CA💼 Full-time💰 $245,000–$245,000🗓 2026-08-03 → 2026-09-26

Core

Design and deploy mental health guardrails for AI interactions, including detection systems, review queues, and interventions connecting users to crisis resources.

Role type

Safeguards Analyst (User Well-being)

Builds

Automated detection models, in-product crisis referral features, and policy enforcement workflows.

Domain

AI Safety / Mental Health / Trust & Safety

Deliverable

production ML models | product features | dashboards & analysis

Required skills

Trust & safety/content moderation experience, mental health harm area expertise, experimental design & evaluation, policy-to-metrics translation, SQL/data analysis, generative AI product experience, risk identification, judgment in ambiguous cases

Preferred skills

Clinical mental health expertise, LLM-based classification system evaluation, agentic tool usage (e.g., Claude Code), crisis support experience

Technologies

SQL, generative AI models, agentic tools

Responsibilities

Support design and execution of interventions with metric definition and dataset curation; Partner with Engineering/Data Science to build and tune detection models; Monitor intervention and detection system performance; Review flagged content for enforcement and policy improvement; Develop in-product crisis resource features; Provide feedback on policy gaps based on real scenarios; Research emerging AI policy and mental health research.

Seniority

Mid-Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.