CareerPlanSign in

Staff Software Development Test Engineer

Bangalore HQ💼 Full-time🗓 2026-07-17 → 2026-09-26

Core

Build automated testing suites and evaluation frameworks for Generative AI, LLM, RAG, and multi-agent systems to detect hallucinations, bias, and data quality failures.

Role type

Staff Software Development Test Engineer (GenAI/LLM Evaluation)

Builds

Automated evaluation pipelines, test beds for multi-agent workflows, and CI/CD quality gates for AI products.

Domain

Automotive retail technology, Generative AI, Machine Learning

Deliverable

production ML models | product features

Required skills

Python (expert), SQL, LLM evaluation frameworks (RAGAS, TruLens, DeepEval), Agent workflow testing (LangChain, LangSmith), Vector DB validation, Statistical analysis (Pandas, NumPy), Test automation (Pytest), API testing (Postman, REST Assured), MLOps (MLflow), CI/CD (GitHub Actions, Jenkins), Observability (Grafana, Kibana, OpenTelemetry)

Preferred skills

Cloud AI Services (AWS Bedrock, Azure OpenAI, GCP Vertex AI), MLOps Platforms (Kubeflow, Weights & Biases), ML Frameworks (Scikit-learn, TensorFlow, PyTorch), Infrastructure as Code (Terraform, Kubernetes), UI Automation (Playwright, Cypress), Performance Engineering (Locust, JMeter), Synthetic Data Generation, Statistical Hypothesis Testing

Technologies

Python, SQL, RAGAS, TruLens, DeepEval, LangChain, LangSmith, LlamaIndex, OpenAI API, Anthropic API, Hugging Face API, Vector DBs, Pandas, NumPy, Pytest, Postman, REST Assured, Requests, MLflow, Docker, GitHub Actions, Jenkins, Grafana, Kibana, OpenTelemetry, AWS Bedrock, Azure OpenAI, GCP Vertex AI, Kubeflow, Weights & Biases, Feast, Scikit-learn, TensorFlow, PyTorch, Terraform, Kubernetes, Playwright, Cypress, Locust, JMeter

Responsibilities

Build automated testing suites to detect hallucinations, bias, toxicity, and prompt injection vulnerabilities; Implement automated evaluations for RAG systems measuring context relevance, groundedness, and answer faithfulness; Design test beds to validate multi-agent workflows including tool-calling accuracy and autonomous decision loops; Build and run automated conversation simulations to stress-test agent behaviour; Create prompt regression frameworks to assess output consistency; Statistically validate AI data outputs and audit data ingestion pipelines; Maintain automated suites tracking ML metrics and deep learning loss curves; Implement continuous monitoring scripts to detect data and concept drift; Build and maintain scalable test automation frameworks for APIs, backend services, and model endpoints; Embed AI evaluation and data QA suites into MLOps and CI/CD pipelines; Define and track AI quality KPIs and communicate release readiness.

Seniority

Staff, hands-on IC with strategic oversight

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.