AI EVALUATION ENGINEER (LLM/Langfuse)
Core
Design and execute evaluation frameworks for LLM outputs, focusing on correctness, hallucination detection, latency, and safety in production environments.
Role type
AI Evaluation Engineer (LLM/Langfuse)
Builds
Robust evaluation pipelines, observability layers, and automated evaluation harnesses integrated with CI/CD.
Domain
Artificial Intelligence / Large Language Models / Observability
Deliverable
production ML models
Required skills
Langfuse, Python, LLM pipelines, tracing, dataset management, scoring pipelines, test harness development, hallucination detection, latency analysis, safety evaluation, LangChain, LangGraph
Preferred skills
RAGAS, DeepEval, HIPAA compliance environments, MCP (Model Context Protocol), LangSmith, product startup experience
Technologies
Langfuse, LangChain, LangGraph, Python, Magpie, RAGAS, DeepEval, LangSmith
Responsibilities
Design and execute evaluation frameworks for LLM outputs; Own Langfuse instrumentation end-to-end including tracing, prompt versioning, dataset management, and scoring pipelines; Build and maintain automated evaluation harnesses integrated with CI/CD pipelines; Collaborate with Forward Deployed Engineers to define success metrics and evaluation criteria; Contribute to Magpie's observability layer including model quality scoring, audit trails, and compliance reporting.
Seniority
Mid-level, hands-on IC