CareerPlanGet AI match score →

Senior Software Development Test Enigneer

Bangalore HQ💼 Full-time🗓 2026-07-17 → 2026-07-31

Core

Build automated testing suites, evaluations, and validation frameworks for Generative AI, LLM, RAG, and multi-agent systems to ensure quality, safety, and reliability before production.

Role type

Senior SDET specializing in GenAI/LLM evaluation and MLOps

Builds

Automated test suites, evaluation frameworks, and CI/CD pipelines for AI/ML products

Domain

Automotive retail technology, Generative AI, Large Language Models (LLM), and Machine Learning Operations (MLOps)

Deliverable

production ML models | product features

Required skills

Python (expert), SQL, LLM evaluation frameworks (RAGAS, TruLens, DeepEval), Agent workflow testing (LangChain, LangSmith), API testing (Postman, REST Assured), CI/CD automation (GitHub Actions, Jenkins), Data quality analysis (Pandas, NumPy), Observability (Grafana, Kibana, OpenTelemetry)

Preferred skills

Cloud AI services (AWS Bedrock, Azure OpenAI, GCP Vertex AI), MLOps platforms (MLflow, Kubeflow, Weights & Biases), ML frameworks (Scikit-learn, TensorFlow, PyTorch), Infrastructure as Code (Docker, Kubernetes, Terraform), UI automation (Playwright, Cypress), Performance engineering (Locust, JMeter), Statistical hypothesis testing, Synthetic data generation

Technologies

Python, SQL, RAGAS, TruLens, DeepEval, Promptflow, LangChain, LangSmith, LlamaIndex, OpenAI API, Anthropic API, Hugging Face API, Vector DBs, Pandas, NumPy, Pytest, Postman, REST Assured, Requests, MLflow, Docker, GitHub Actions, Jenkins, Grafana, Kibana, OpenTelemetry, AWS Bedrock, Azure OpenAI, GCP Vertex AI, MLflow, Kubeflow, Weights & Biases, Feast, Scikit-learn, TensorFlow, PyTorch, Docker, Kubernetes, Terraform, Playwright, Cypress, Locust, JMeter

Responsibilities

Build automated testing suites to detect hallucinations, bias, toxicity, and prompt injection vulnerabilities; Implement automated evaluations for RAG systems measuring context relevance, groundedness, and answer faithfulness; Design test beds to validate multi-agent workflows; Build and run automated conversation simulations; Create prompt regression frameworks; Statistically validate AI data outputs; Programmatically audit data ingestion and transformation pipelines; Validate vector DB indexing and retrieval latency; Maintain automated suites tracking ML metrics and deep learning loss curves; Implement continuous monitoring scripts to detect data and concept drift; Build and maintain scalable test automation frameworks for APIs, backend services, and model endpoints; Embed AI evaluation and data QA suites into MLOps and CI/CD pipelines

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗