Senior AI Backend Engineer - Agent Evaluation & Quality
Core
Building evaluation systems (judges, test harnesses, simulators) to measure multi-agent system performance, catch regressions, and ensure quality before release.
Role type
Senior IC backend engineer specializing in LLM/agent evaluation infrastructure
Builds
Evaluation stack, LLM-as-judge systems, CI/CD regression harnesses, user simulators, and production failure datasets
Domain
AI/LLM multi-agent systems, software engineering infrastructure
Deliverable
production ML models | infrastructure
Required skills
Python or TypeScript, system design, CI/CD integration, LLM/agent development (RAG, tool calling, orchestration), metrics calibration, production reliability engineering
Preferred skills
LLM-as-judge evaluation, observability tooling (Arize, LangSmith), Arabic NLP, e-commerce product experience
Technologies
Python, TypeScript, LangGraph, LangChain, CI/CD pipelines
Responsibilities
Design and build LLM-as-judge systems calibrated against human labels; build per-PR eval harnesses and regression detection for CI; create user simulators for adversarial test coverage; pipe production failures into evaluation sets; define measurable quality criteria with product teams; contribute to agent development based on evaluation insights
Seniority
Senior, hands-on IC
