LLM Evaluation Engineer
Core
Building the evaluation layer for an AI control platform that monitors, evaluates, and enforces safety policies on LLM prompts, responses, and agent behaviors in real-time.
Role type
Senior IC LLM Evaluation Engineer
Builds
Real-time guardrails, classifiers, semantic judgment systems, and policy enforcement logic for enterprise AI.
Domain
Enterprise AI Safety / LLM Observability
Deliverable
production ML models
Required skills
LLM systems engineering, foundation model integration, vector search, semantic similarity, rule-based systems, Python, prompt engineering, model debugging, production deployment
Preferred skills
OpenTelemetry, Model Context Protocol (MCP), red-teaming, AI risk taxonomies
Technologies
OpenAI, Claude, Mistral, Llama, FAISS, Qdrant, Weaviate, Hugging Face Transformers, LangChain, PyTorch, TensorFlow
Responsibilities
Design real-time evaluation logic for policy violations; Implement evaluation strategies using semantic similarity and classifiers; Integrate model outputs with enforcement actions; Prototype and tune small language models; Collaborate on data infrastructure connections; Build debugging and observation tools; Define reusable evaluation abstractions
Seniority
Senior, hands-on IC