AI Engineer (Harness)
Core
Design, evaluate, and benchmark infrastructure for AI models and autonomous agentic workloads, serving as the bridge between data science and production agent scaffolding.
Role type
Senior IC AI Evaluation Engineer (Agentic Systems)
Builds
Scalable evaluation harnesses, LLM-as-a-judge pipelines, and benchmark datasets for multi-step agent systems.
Domain
Artificial Intelligence / Autonomous Agents / Statistical Validation
Deliverable
production ML models
Required skills
Statistical experimental design, LLM-as-a-judge validation, evaluation dataset engineering, non-deterministic system evaluation, metric design, error analysis, Python, LLM integration, agent architecture awareness
Preferred skills
Agent frameworks (LangChain, LangGraph, AutoGen), secure enclave/Kubernetes environments
Technologies
Python, LLM APIs, LangChain, LangGraph, AutoGen, lm-evaluation-harness, Promptfoo, Ragas
Responsibilities
Architect scalable Python evaluation harnesses for multi-step agentic systems; Apply statistical methods (hypothesis testing, bootstrapping) to validate performance changes; Build and calibrate LLM-as-a-judge pipelines against human ground truth; Curate benchmark datasets and scoring rubrics; Conduct deep error and trajectory analysis on agent runs; Translate evaluation results into architectural improvements for agent scaffolding and safety guardrails
Seniority
Senior, hands-on IC
