LLM Engineer (LLM Evaluation)
Core
Design and build evaluation systems and platforms to reliably assess Large Language Model (LLM) performance and continuously improve model quality.
Role type
LLM Evaluation Engineer
Builds
Benchmark datasets, evaluation protocols, and automation pipelines for LLM quality validation
Domain
Artificial Intelligence / Large Language Models / MLOps
Deliverable
production ML models
Required skills
LLM evaluation framework design, benchmark dataset construction, evaluation metric design (Human/LLM-based), evaluation protocol establishment, reproducibility assurance, evaluation automation, regression detection, model quality validation
Preferred skills
Kubernetes development, large-scale data processing, GPU-based distributed inference, monitoring tooling (Datadog, Prometheus), MLflow/Argo Workflows operation, GPU cluster pipeline design
Technologies
Argo Workflows, MLflow, Python, Kubernetes, Datadog, Prometheus
Responsibilities
Design LLM evaluation benchmarks and metrics; Establish fair evaluation protocols and ensure reproducibility; Build evaluation automation environments and integrate ML pipelines; Design automated regression detection and alerting systems; Operate continuous model quality improvement processes based on evaluation results
Seniority
Mid-Senior, hands-on IC