CareerPlanSign in

Manager, AI Benchmarking and Evaluation Research (Remote, ROU)

Romania - Remote🌐 Remote💼 Full-time🗓 2026-09-08 → 2026-09-25

Core

Lead a team to design and build rigorous evaluation standards, datasets, and methodologies for AI/LLM and agentic systems operating in cybersecurity workflows.

Role type

Senior Manager, AI Evaluation Research (Cybersecurity)

Builds

Benchmark datasets, reproducible testing pipelines, and evaluation frameworks for AI security models

Domain

Cybersecurity / AI Safety & Evaluation

Deliverable

production ML models

Required skills

SOC operations expertise, team leadership, incident response, threat hunting, AI/LLM knowledge, evaluation methodology design, cross-functional collaboration

Preferred skills

LLM/agentic evaluation frameworks, Python programming, SIEM/SOAR tools, adversarial testing, MITRE ATT&CK

Technologies

LLMs, agentic systems, SIEM, SOAR, Python

Responsibilities

Lead and mentor researchers and engineers; define strategy and success metrics for AI evaluation; design standardized testing pipelines; assess model accuracy and robustness; translate findings into engineering improvements; communicate results to technical and executive audiences

Seniority

Manager, hands-on leadership

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.