CareerPlanSign in

AI Engineer (Harness)

London💼 Full-time🗓 2026-09-16 → 2026-09-26

Core

Design, evaluate, and benchmark infrastructure for AI models and autonomous agentic workloads, serving as the bridge between data science and production agent scaffolding.

Role type

Senior IC AI Evaluation Engineer (Agentic Systems)

Builds

Scalable evaluation harnesses, LLM-as-a-judge pipelines, and benchmark datasets for multi-step agent systems.

Domain

Artificial Intelligence / Autonomous Agents / Statistical Validation

Deliverable

production ML models

Required skills

Statistical experimental design, LLM-as-a-judge validation, evaluation dataset engineering, non-deterministic system evaluation, metric design, error analysis, Python, LLM integration, agent architecture awareness

Preferred skills

Agent frameworks (LangChain, LangGraph, AutoGen), secure enclave/Kubernetes environments

Technologies

Python, LLM APIs, LangChain, LangGraph, AutoGen, lm-evaluation-harness, Promptfoo, Ragas

Responsibilities

Architect scalable Python evaluation harnesses for multi-step agentic systems; Apply statistical methods (hypothesis testing, bootstrapping) to validate performance changes; Build and calibrate LLM-as-a-judge pipelines against human ground truth; Curate benchmark datasets and scoring rubrics; Conduct deep error and trajectory analysis on agent runs; Translate evaluation results into architectural improvements for agent scaffolding and safety guardrails

Seniority

Senior, hands-on IC

Sourced via lever · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.