Senior Security Researcher and Principal Security Researcher (Multiple Positions)
Core
Design and build end-to-end evaluation harnesses for agentic security and engineering workflows, creating benchmark suites and automated grading pipelines to measure model safety, correctness, and efficiency.
Role type
Senior/Principal Security Researcher (AI Agent Evaluation)
Builds
Evaluation harnesses, benchmark datasets, automated graders, and integration into CI/CD release gates.
Domain
AI Security / Agentic Workflows / Model Evaluation
Deliverable
production ML models | product features
Required skills
software engineering, system design, LLMs and AI agents, model evaluation, prompt orchestration, security engineering, vulnerability management, statistical analysis, automation workflows, structured data processing, regression analysis, human-review protocols
Preferred skills
experience with SAST/SCA, SARIF, secure development lifecycle, compliance-sensitive systems, experimentation design, quality measurement systems
Technologies
LLMs, AI agents, SAST, SCA, SARIF, CI/CD, APIs, automation frameworks
Responsibilities
Design and implement evaluation systems end-to-end including task definition, dataset creation, harness implementation, scoring, and operational integration. Create representative benchmark suites and golden datasets covering normal, edge, adversarial, and production-derived cases. Integrate evaluations into engineering workflows and release gates to support reproducible evidence before production rollout. Detect regressions, benchmark drift, and cases where eval scores improve while real-world outcomes do not. Partner with agent builders and product teams to turn evaluation results into release recommendations and autonomy-boundary decisions. Convert learnings into reusable paved paths including playbooks, templates, dashboards, and reference implementations.
Seniority
Senior/Principal, hands-on IC with strategy & mentorship