Senior Software Development Test Enigneer
Core
Build automated testing suites, evaluations, and validation frameworks for Generative AI, LLM, RAG, and multi-agent systems to ensure quality, safety, and reliability before production.
Role type
Senior SDET specializing in GenAI/LLM evaluation and MLOps
Builds
Automated test suites, evaluation frameworks, and CI/CD pipelines for AI/ML products
Domain
Automotive retail technology, Generative AI, Large Language Models (LLM), and Machine Learning Operations (MLOps)
Deliverable
production ML models | product features
Required skills
Python (expert), SQL, LLM evaluation frameworks (RAGAS, TruLens, DeepEval), Agent workflow testing (LangChain, LangSmith), API testing (Postman, REST Assured), CI/CD automation (GitHub Actions, Jenkins), Data quality analysis (Pandas, NumPy), Observability (Grafana, Kibana, OpenTelemetry)
Preferred skills
Cloud AI services (AWS Bedrock, Azure OpenAI, GCP Vertex AI), MLOps platforms (MLflow, Kubeflow, Weights & Biases), ML frameworks (Scikit-learn, TensorFlow, PyTorch), Infrastructure as Code (Docker, Kubernetes, Terraform), UI automation (Playwright, Cypress), Performance engineering (Locust, JMeter), Statistical hypothesis testing, Synthetic data generation
Technologies
Python, SQL, RAGAS, TruLens, DeepEval, Promptflow, LangChain, LangSmith, LlamaIndex, OpenAI API, Anthropic API, Hugging Face API, Vector DBs, Pandas, NumPy, Pytest, Postman, REST Assured, Requests, MLflow, Docker, GitHub Actions, Jenkins, Grafana, Kibana, OpenTelemetry, AWS Bedrock, Azure OpenAI, GCP Vertex AI, MLflow, Kubeflow, Weights & Biases, Feast, Scikit-learn, TensorFlow, PyTorch, Docker, Kubernetes, Terraform, Playwright, Cypress, Locust, JMeter
Responsibilities
Build automated testing suites to detect hallucinations, bias, toxicity, and prompt injection vulnerabilities; Implement automated evaluations for RAG systems measuring context relevance, groundedness, and answer faithfulness; Design test beds to validate multi-agent workflows; Build and run automated conversation simulations; Create prompt regression frameworks; Statistically validate AI data outputs; Programmatically audit data ingestion and transformation pipelines; Validate vector DB indexing and retrieval latency; Maintain automated suites tracking ML metrics and deep learning loss curves; Implement continuous monitoring scripts to detect data and concept drift; Build and maintain scalable test automation frameworks for APIs, backend services, and model endpoints; Embed AI evaluation and data QA suites into MLOps and CI/CD pipelines
Seniority
Senior, hands-on IC