Senior Research Engineer, Safety
Core
Building safeguards, evaluations, and runtime controls to ensure AI agents are safe, reliable, and controllable in production environments.
Role type
Senior Research Engineer (AI Safety)
Builds
Production-ready safety models, adversarial evaluation suites, red-team datasets, and runtime safeguards for conversational AI agents.
Domain
Applied Conversational AI / Enterprise Safety
Deliverable
production ML models
Required skills
AI/ML engineering, model evaluation, post-training techniques (RL, preference optimization, distillation), adversarial testing, prompt injection mitigation, policy enforcement, Python, experimental judgment
Preferred skills
High-stakes enterprise workflow safeguards, human-in-the-loop review, incident response frameworks
Technologies
Python, Reinforcement Learning, Preference Optimization, Distillation, Synthetic Data Generation
Responsibilities
Research and build safeguards against prompt injection, unsafe tool use, and hallucinated commitments; Build adversarial evaluations and red-team datasets; Develop classifiers, judges, and reward signals; Analyze production traces to identify root causes and test mitigations; Partner with cross-functional teams to rollout scalable safety practices.
Seniority
Senior, hands-on IC