CareerPlanSign in

AI EVALUATION ENGINEER (LLM/Langfuse)

🌐 Remote💼 Full-time🗓 2026-06-25

Core

Design and execute evaluation frameworks for LLM outputs, focusing on correctness, hallucination detection, latency, and safety in production environments.

Role type

AI Evaluation Engineer (LLM/Langfuse)

Builds

Robust evaluation pipelines, observability layers, and automated evaluation harnesses integrated with CI/CD.

Domain

Artificial Intelligence / Large Language Models / Observability

Deliverable

production ML models

Required skills

Langfuse, Python, LLM pipelines, tracing, dataset management, scoring pipelines, test harness development, hallucination detection, latency analysis, safety evaluation, LangChain, LangGraph

Preferred skills

RAGAS, DeepEval, HIPAA compliance environments, MCP (Model Context Protocol), LangSmith, product startup experience

Technologies

Langfuse, LangChain, LangGraph, Python, Magpie, RAGAS, DeepEval, LangSmith

Responsibilities

Design and execute evaluation frameworks for LLM outputs; Own Langfuse instrumentation end-to-end including tracing, prompt versioning, dataset management, and scoring pipelines; Build and maintain automated evaluation harnesses integrated with CI/CD pipelines; Collaborate with Forward Deployed Engineers to define success metrics and evaluation criteria; Contribute to Magpie's observability layer including model quality scoring, audit trails, and compliance reporting.

Seniority

Mid-level, hands-on IC

Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.