CareerPlanGet AI match score →

AI EVALUATION ENGINEER (LLM/Langfuse)

🌐 Remote💼 Full-time🗓 2026-06-25

Core

Design and execute evaluation frameworks for LLM outputs, focusing on correctness, hallucination detection, latency, and safety in production environments.

Role type

AI Evaluation Engineer (LLM/Langfuse)

Builds

Robust evaluation pipelines, observability layers, and automated evaluation harnesses integrated with CI/CD.

Domain

Artificial Intelligence / Large Language Models / Observability

Deliverable

production ML models

Required skills

Langfuse, Python, LLM pipelines, tracing, dataset management, scoring pipelines, test harness development, hallucination detection, latency analysis, safety evaluation, LangChain, LangGraph

Preferred skills

RAGAS, DeepEval, HIPAA compliance environments, MCP (Model Context Protocol), LangSmith, product startup experience

Technologies

Langfuse, LangChain, LangGraph, Python, Magpie, RAGAS, DeepEval, LangSmith

Responsibilities

Design and execute evaluation frameworks for LLM outputs; Own Langfuse instrumentation end-to-end including tracing, prompt versioning, dataset management, and scoring pipelines; Build and maintain automated evaluation harnesses integrated with CI/CD pipelines; Collaborate with Forward Deployed Engineers to define success metrics and evaluation criteria; Contribute to Magpie's observability layer including model quality scoring, audit trails, and compliance reporting.

Seniority

Mid-level, hands-on IC

Rewrite
## About the Role Steinn Labs is looking for an AI Evaluation Engineer to own the quality, reliability, and observability of LLM-powered systems across real client deployments. This is a hands-on role focused on evaluating, instrumenting, and improving AI agent outputs in production environments. You will work closely with our Forward Deployed Engineers (FDEs) and contribute to building robust evaluation pipelines and observability layers that ensure our AI systems meet real-world performance and compliance standards. ## Responsibilities - Design and execute evaluation frameworks for LLM outputs, focusing on correctness, hallucination detection, latency, and safety. - Own Langfuse instrumentation end-to-end, including: - Tracing - Prompt versioning - Dataset management - Scoring pipelines - Build and maintain automated evaluation harnesses, integrated with CI/CD pipelines to catch regressions before deployment. - Collaborate with FDEs to define success metrics and evaluation criteria for new agent features and Proof Sprint deliverables. - Contribute to Magpie's observability layer, including: - Model quality scoring - Audit trails - Compliance reporting (HIPAA, SOC 2) ## Requirements - Langfuse (Hard Requirement) - Proven hands-on experience instrumenting production pipelines - Experience with tracing, datasets, scoring pipelines (not just tutorials) - LangChain / LangGraph - Experience building or working with LLM pipelines - Python - Strong experience in scripting, evaluation logic, and test harness development - LLM Evaluation Expertise - Strong judgment on output quality, hallucinations, and response reliability - Minimum 2+ years in AI/ML or LLM-focused engineering roles ## Nice to Have - Experience with evaluation frameworks like RAGAS, DeepEval, or similar - Exposure to healthcare data / HIPAA compliance environments - Experience with MCP (Model Context Protocol) or tool-call tracing - Familiarity with LangSmith or similar observability tools - Background in product startups, AI labs, or product engineering agencies ## Key Competencies - Strong analytical thinking and attention to detail - Ability to evaluate AI outputs beyond surface-level correctness - Ownership mindset — ability to independently drive quality improvements - Strong collaboration skills across engineering and client-facing teams ## What We Offer - Work on real-world AI systems deployed for US enterprise clients - Own critical parts of the AI quality and observability stack - Fast-paced environment with high ownership and growth opportunities - Exposure to cutting-edge agentic AI systems and evaluation frameworks ## About the Company Steinn Labs Pvt. Ltd. ## How to Apply - Share your resume at: [email protected]
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗