Senior Software Development Test Enigneer
Core
Build automated testing suites and evaluation frameworks for Generative AI, LLM, RAG, and multi-agent systems to detect hallucinations, bias, and data quality issues.
Role type
Senior SDET specializing in GenAI/LLM evaluation and MLOps
Builds
Automated test suites, evaluation pipelines, and CI/CD quality gates for AI products
Domain
Generative AI, Large Language Models (LLM), RAG systems, and MLOps
Deliverable
production ML models | product features
Required skills
Python (expert), SQL, Pytest, RAGAS/TruLens/DeepEval, LangChain, Vector DB testing, Statistical analysis
Preferred skills
AWS Bedrock/Azure OpenAI/GCP Vertex AI, MLflow, Kubernetes, Playwright, Locust/JMeter, Synthetic data generation
Technologies
Python, SQL, Pytest, RAGAS, TruLens, DeepEval, LangChain, LangSmith, LlamaIndex, OpenAI API, Anthropic API, Hugging Face, Vector DBs, Pandas, NumPy, MLflow, Docker, GitHub Actions, Jenkins, Grafana, Kibana, OpenTelemetry
Responsibilities
Build automated testing suites to detect hallucinations, bias, and prompt injection vulnerabilities; Implement automated evaluations for RAG systems; Design test beds to validate multi-agent workflows; Build and run automated conversation simulations; Statistically validate AI data outputs; Maintain automated suites tracking ML metrics; Build and maintain scalable test automation frameworks for APIs and model endpoints
Seniority
Senior, hands-on IC