CareerPlanGet AI match score →

Staff+ Software Engineer, Safeguards Review Tooling

San Francisco, CA💼 Full-time💰 $320,000–$320,000🗓 2026-07-13 → 2026-07-31

Core

Building internal investigation, review, and enforcement tooling for Anthropic's AI safety team to help humans and AI agents identify and act on harmful behavior across first-party and third-party platforms.

Role type

Staff+ Software Engineer (Safety/Trust & Safety Tooling)

Builds

Case queues, investigation views, decision logging, account-actioning workflows, and the underlying platform APIs/data storage for safety review.

Domain

AI Safety / Trust & Safety / Platform Engineering

Deliverable

production ML models | product features | infrastructure

Required skills

Full-stack or platform engineering, architecture and design, shipping internal tools for demanding operational users, cross-functional collaboration with non-engineering teams

Preferred skills

Trust and safety/integrity/fraud/abuse-prevention tooling, designing systems under strict privacy/compliance constraints, integrating LLMs/agentic systems into operational workflows, building developer platforms/extensible tooling frameworks, supporting enforcement systems across multiple product surfaces

Technologies

None explicitly stated

Responsibilities

Build investigation and enforcement tooling including case queues and audit logging; Develop platform layer of reusable APIs and backend services; Scale review through automation including Claude-assisted workflows; Partner with policy, legal, and privacy stakeholders to translate needs into reliable systems; Build guardrails including granular permissions and audit trails; Instrument tools with metrics on queue health and decision quality

Seniority

Staff+, hands-on IC with architectural ownership

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗