Staff Software Development Test Engineer
Core
Build automated testing suites and evaluation frameworks for Generative AI, LLM, RAG, and multi-agent systems to detect hallucinations, bias, and data quality failures.
Role type
Staff Software Development Test Engineer (GenAI/LLM Evaluation)
Builds
Automated evaluation pipelines, test beds for multi-agent workflows, and CI/CD quality gates for AI products.
Domain
Automotive retail technology, Generative AI, Machine Learning
Deliverable
production ML models | product features
Required skills
Python (expert), SQL, LLM evaluation frameworks (RAGAS, TruLens, DeepEval), Agent workflow testing (LangChain, LangSmith), Vector DB validation, Statistical analysis (Pandas, NumPy), Test automation (Pytest), API testing (Postman, REST Assured), MLOps (MLflow), CI/CD (GitHub Actions, Jenkins), Observability (Grafana, Kibana, OpenTelemetry)
Preferred skills
Cloud AI Services (AWS Bedrock, Azure OpenAI, GCP Vertex AI), MLOps Platforms (Kubeflow, Weights & Biases), ML Frameworks (Scikit-learn, TensorFlow, PyTorch), Infrastructure as Code (Terraform, Kubernetes), UI Automation (Playwright, Cypress), Performance Engineering (Locust, JMeter), Synthetic Data Generation, Statistical Hypothesis Testing
Technologies
Python, SQL, RAGAS, TruLens, DeepEval, LangChain, LangSmith, LlamaIndex, OpenAI API, Anthropic API, Hugging Face API, Vector DBs, Pandas, NumPy, Pytest, Postman, REST Assured, Requests, MLflow, Docker, GitHub Actions, Jenkins, Grafana, Kibana, OpenTelemetry, AWS Bedrock, Azure OpenAI, GCP Vertex AI, Kubeflow, Weights & Biases, Feast, Scikit-learn, TensorFlow, PyTorch, Terraform, Kubernetes, Playwright, Cypress, Locust, JMeter
Responsibilities
Build automated testing suites to detect hallucinations, bias, toxicity, and prompt injection vulnerabilities; Implement automated evaluations for RAG systems measuring context relevance, groundedness, and answer faithfulness; Design test beds to validate multi-agent workflows including tool-calling accuracy and autonomous decision loops; Build and run automated conversation simulations to stress-test agent behaviour; Create prompt regression frameworks to assess output consistency; Statistically validate AI data outputs and audit data ingestion pipelines; Maintain automated suites tracking ML metrics and deep learning loss curves; Implement continuous monitoring scripts to detect data and concept drift; Build and maintain scalable test automation frameworks for APIs, backend services, and model endpoints; Embed AI evaluation and data QA suites into MLOps and CI/CD pipelines; Define and track AI quality KPIs and communicate release readiness.
Seniority
Staff, hands-on IC with strategic oversight