LLM Engineer (LLM Evaluation)
Core
Designing and operating evaluation systems and platforms to reliably assess Large Language Model (LLM) performance and continuously improve model quality.
Role type
LLM Evaluation Engineer
Builds
Benchmark datasets, evaluation protocols, automation pipelines, and end-to-end evaluation workflows for LLMs.
Domain
Artificial Intelligence / Large Language Models / Automotive Software
Deliverable
production ML models
Required skills
LLM evaluation framework design, benchmark dataset construction, evaluation metric design (Human/LLM-based), evaluation protocol establishment, reproducibility assurance, evaluation automation, ML pipeline integration, regression detection, model quality validation
Preferred skills
Kubernetes development, large-scale data processing, distributed inference, monitoring tooling (Datadog, Prometheus), MLflow/Argo Workflows operation, GPU cluster management
Technologies
Python, Argo Workflows, MLflow, Kubernetes, Datadog, Prometheus, lm-eval, HELM, OpenAI Evals
Responsibilities
Design LLM evaluation benchmarks and metrics; Establish fair evaluation protocols and ensure reproducibility; Build and integrate automated evaluation workflows; Design systems for automatic detection of model performance regression; Operate continuous model quality improvement processes
Seniority
Mid-Senior, hands-on IC