CareerPlanSign in

Senior AI Backend Engineer - Agent Evaluation & Quality

Makkah, Makkah Province, Saudi Arabia💼 Full-time🗓 2026-07-19 → 2026-09-27

Core

Building evaluation systems (judges, test harnesses, simulators) to measure multi-agent system performance, catch regressions, and ensure quality before release.

Role type

Senior IC backend engineer specializing in LLM/agent evaluation infrastructure

Builds

Evaluation stack, LLM-as-judge systems, CI/CD regression harnesses, user simulators, and production failure datasets

Domain

AI/LLM multi-agent systems, software engineering infrastructure

Deliverable

production ML models | infrastructure

Required skills

Python or TypeScript, system design, CI/CD integration, LLM/agent development (RAG, tool calling, orchestration), metrics calibration, production reliability engineering

Preferred skills

LLM-as-judge evaluation, observability tooling (Arize, LangSmith), Arabic NLP, e-commerce product experience

Technologies

Python, TypeScript, LangGraph, LangChain, CI/CD pipelines

Responsibilities

Design and build LLM-as-judge systems calibrated against human labels; build per-PR eval harnesses and regression detection for CI; create user simulators for adversarial test coverage; pipe production failures into evaluation sets; define measurable quality criteria with product teams; contribute to agent development based on evaluation insights

Seniority

Senior, hands-on IC

Sourced via workable · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.