Staff+ Software Engineer, Safeguards Review Tooling
Core
Building internal investigation, review, and enforcement tooling for Anthropic's AI safety team to help humans and AI agents identify and act on harmful behavior across first-party and third-party platforms.
Role type
Staff+ Software Engineer (Safety/Trust & Safety Tooling)
Builds
Case queues, investigation views, decision logging, account-actioning workflows, and the underlying platform APIs/data storage for safety review.
Domain
AI Safety / Trust & Safety / Platform Engineering
Deliverable
production ML models | product features | infrastructure
Required skills
Full-stack or platform engineering, architecture and design, shipping internal tools for demanding operational users, cross-functional collaboration with non-engineering teams
Preferred skills
Trust and safety/integrity/fraud/abuse-prevention tooling, designing systems under strict privacy/compliance constraints, integrating LLMs/agentic systems into operational workflows, building developer platforms/extensible tooling frameworks, supporting enforcement systems across multiple product surfaces
Technologies
None explicitly stated
Responsibilities
Build investigation and enforcement tooling including case queues and audit logging; Develop platform layer of reusable APIs and backend services; Scale review through automation including Claude-assisted workflows; Partner with policy, legal, and privacy stakeholders to translate needs into reliable systems; Build guardrails including granular permissions and audit trails; Instrument tools with metrics on queue health and decision quality
Seniority
Staff+, hands-on IC with architectural ownership