Staff Software Development Test Engineer - AI Evaluation
Core
Building Tekion's shared AI evaluation platform and frameworks to measure, trust, and improve the quality of AI outputs across ML teams.
Role type
Senior Staff SDET specializing in AI Evaluation and ML Quality Engineering
Builds
Automated scoring pipelines, evaluation datasets, quality metrics, and CI/CD evaluation gates for AI agents and models
Domain
Automotive retail technology, Generative AI, Machine Learning
Deliverable
production ML models | dashboards & analysis
Required skills
Python, LLM evaluation, benchmark design, dataset curation, statistical analysis, automated scoring pipelines, LLM-as-judge, RAG concepts, hallucination detection, CI/CD integration
Preferred skills
Eval frameworks (Ragas, DeepEval, LangSmith, TruLens, Promptfoo, HELM), online evaluation, guardrails, red-teaming, experiment tracking (MLflow, Weights & Biases), A/B testing
Technologies
Python, LLMs, Ragas, DeepEval, LangSmith, TruLens, Promptfoo, HELM, MLflow, Weights & Biases
Responsibilities
Design and own AI evaluation infrastructure; create and maintain evaluation datasets and ground-truth sets; define quality metrics for AI outputs; build automated scoring pipelines including LLM-as-judge; validate user intents and measure response accuracy; identify hallucinations and edge cases; establish evaluation gates in CI/CD; develop dashboards for AI quality visibility; use AI/LLMs to scale evaluation via automated judges and synthetic data
Seniority
Senior, hands-on IC