CareerPlanGet AI match score →

Research Engineer, Safeguards Labs

New York City, NY💼 Full-time💰 $350,000–$350,000🗓 2026-05-21 → 2026-07-31

Core

Define and execute research agenda to develop novel safety methods, detect misuse of Claude, and build classifiers for abuse patterns.

Role type

Research Engineer (AI Safety & Safeguards)

Builds

Offline analysis systems, detection classifiers, and prototypes for real-time safeguards.

Domain

Artificial Intelligence, Large Language Models, AI Safety, Trust & Safety

Deliverable

production ML models

Required skills

Python, large dataset analysis, independent project scoping, LLM operation familiarity (sampling/prompting/training), experimental design

Preferred skills

Machine learning model training, evaluation methodologies for language models, agentic environment evaluation, red teaming/jailbreak research, interpretability methods, prototype-to-production transfer

Technologies

Python

Responsibilities

Lead research projects on detecting misuse and malicious accounts; Design and run offline analyses over model usage data; Develop and iterate on prototypes for real-time safeguards; Contribute to research on detecting abusive behavior in chat/agentive workflows; Build evaluations for measuring safeguard effectiveness; Write findings to inform decisions across teams.

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗