LLMOps Engineer
Core
Build and run evaluation suites for an AI learning platform's LLM products to verify functionality and detect model failures before release.
Role type
Junior-to-mid LLMOps Engineer (Evaluation & Measurement)
Builds
Evaluation suites, reporting dashboards, and core datasets for LLM products
Domain
EdTech / Generative AI / LLM Operations
Deliverable
production ML models | dashboards & analysis
Required skills
Python, data analysis, LLM evaluation decomposition, familiarity with Langfuse/RAGAS/DSPy
Preferred skills
cost optimization at scale, dashboarding, classical statistics
Technologies
Langfuse, RAGAS, DSPy
Responsibilities
Build and run eval suites per product decomposed by failure mode; maintain cost-effective evals as product line grows; own reporting loop for findings and platform data; steward core datasets including classified customer-support data; partner with AI Product Manager on instrumentation for evals
Seniority
Junior-to-mid, hands-on IC with growth path