Manager, AI Benchmarking and Evaluation Research (Remote, ROU)
Core
Lead a team to design and build rigorous evaluation standards, datasets, and methodologies for AI/LLM and agentic systems operating in cybersecurity workflows.
Role type
Senior Manager, AI Evaluation Research (Cybersecurity)
Builds
Benchmark datasets, reproducible testing pipelines, and evaluation frameworks for AI security models
Domain
Cybersecurity / AI Safety & Evaluation
Deliverable
production ML models
Required skills
SOC operations expertise, team leadership, incident response, threat hunting, AI/LLM knowledge, evaluation methodology design, cross-functional collaboration
Preferred skills
LLM/agentic evaluation frameworks, Python programming, SIEM/SOAR tools, adversarial testing, MITRE ATT&CK
Technologies
LLMs, agentic systems, SIEM, SOAR, Python
Responsibilities
Lead and mentor researchers and engineers; define strategy and success metrics for AI evaluation; design standardized testing pipelines; assess model accuracy and robustness; translate findings into engineering improvements; communicate results to technical and executive audiences
Seniority
Manager, hands-on leadership