Research Engineer - Agent Intelligence & Evaluation
Core
Building evaluation infrastructure, observability, and self-improvement systems for self-healing voice agents in enterprise customer support.
Role type
Senior Research Engineer (Agent Intelligence & Evaluation)
Builds
Self-healing voice agents for enterprise customer support
Domain
Voice AI, LLM agents, enterprise customer support
Deliverable
production ML models
Required skills
Python, modern ML tooling, speech and audio models, LLM agent systems, evaluation infrastructure, adversarial dataset generation, LLM-as-judge rubrics, pipeline tracing, failure pattern mining, prompt engineering, adversarial replay, regression guardrails
Preferred skills
real-time systems, telephony, RLHF, DPO, synthetic data pipelines, enterprise deployment (SOC 2, PII, data residency)
Technologies
Python, LLM frameworks, audio processing libraries, tracing/observability tools
Responsibilities
Design and maintain evaluation infrastructure with audio-native metrics; implement observability across the pipeline to correlate audio, STT, LLM reasoning, tool calls, and TTS; build self-improvement systems to mine production traces and generate training data; validate fixes with adversarial replay and guardrail against regressions
Seniority
Senior, hands-on IC