CareerPlanGet AI match score →

QA Engineer, AI Quality and Automation - HeartStamp

🌐 Remote💼 Full-time💰 $1,040–$1,300🗓 2026-06-25

Core

Stabilize and extend Playwright end-to-end test suites for AI-generated greeting cards and customer storefronts; build a structured evaluation system to score and improve AI output quality.

Role type

QA Engineer, AI Quality and Automation

Builds

Production-ready test suites and automated evaluation pipelines for AI-generated content

Domain

Generative AI / E-commerce / Print-on-demand

Deliverable

production ML models | product features

Required skills

Playwright automation, LLM evaluation rubric design, regression tracking, full-stack web testing (Next.js/React, Python/FastAPI), RAG and AI orchestration systems understanding

Preferred skills

LangChain, LangGraph, LangSmith, AI-assisted testing tools, prompt-based test generation

Technologies

Playwright, LangChain, LangGraph, LangSmith, Linear, Slack, Google Workspace, Sentry, Grafana, Hotjar, pgvector, LangGraph

Responsibilities

Audit and stabilize existing end-to-end test suites; extend coverage for AI generation and storefront surfaces; design and implement rubrics for AI output quality; translate evaluation findings into actionable items for engineering teams; produce concise async reports on test and eval status

Seniority

Mid-level (3-5 years experience), hands-on IC

Rewrite
## About the role The mandate is narrow by design. You will inherit our Playwright end-to-end suite, stabilize it, and extend coverage where it matters most. Then you will build out a structured eval system that scores and continuously improves AI output quality across the platform. At 20 to 25 hours a week, scope discipline is the job. We want someone who finds the highest-signal coverage gaps, closes them with precision, and builds rubrics that tell the team what good output actually looks like. Some people do their best work in a tightly scoped role with real accountability. If that is you, keep reading. This can scale to full-time for the right person. ## What you walk into The Playwright suite covers core happy paths across Stampy and the storefront. Some tests are stable, a handful are flaky and known. The LangChain eval setup for Stampy is early but functional. Real ground to build on, real cleanup to do. ## What you will own - Playwright automation. Take over the existing e2e suite, audit it, stabilize what is flaky, then extend coverage on the two surfaces that matter most: the Stampy generation experience and the customer storefront. - LLM eval system. Build defensible rubrics for AI card output covering prompt fidelity, aesthetic quality, text accuracy, and style consistency. Inherit and extend our LangChain eval setup to score against those rubrics and flag regressions before they reach users. - Close the loop. Translate eval findings into clear, actionable items for the AI engineer and product team. - Reporting. Tight async updates the team can read in two minutes and know exactly what needs attention. ## First 30 / 60 / 90 - Day 30: full suite reviewed, biggest gaps found, first round of new or stabilized tests shipped, first-draft eval rubric run against real output. - Day 60: evals running on a set cadence, rubric reviewed with the AI engineer and holding up, at least two regressions surfaced and fixed with traceable findings. - Day 90: coverage visibly better on both surfaces, eval scores trending up, reports short and trusted. ## Requirements - 3 to 5 years in QA engineering with real ownership of an automation framework, not just contributing to one. - Strong Playwright. You can read an unfamiliar suite, understand its structure, and extend it confidently. - Hands-on experience evaluating visual or multimodal AI output. Our product is greeting cards and the rubrics score aesthetics, style, and composition. This is core. - Experience building and running LLM eval pipelines: rubric design, scoring logic, regression tracking. - Enough full stack literacy to test against a modern web app (Next.js / React front end, Python / FastAPI back end). - Working understanding of RAG and AI orchestration systems, enough to write meaningful tests and evals against them. We run pgvector, LangChain, and LangGraph. - Clear written communication. Your updates are specific enough that engineers act without a follow-up. - Comfortable working independently in a fast-moving remote team. Consistent EST overlap and fluent C-1 English required. ## Nice to have - LangChain, LangGraph, or LangSmith experience. - AI-assisted testing tools or prompt-based test generation. - E-commerce, print-on-demand, or consumer product background. - Evals that directly informed prompt engineering or model decisions. ## Stack - Playwright - LangChain eval framework - Custom rubrics - LLM-as-judge patterns - Linear - Slack - Google Workspace - Sentry - Grafana - Hotjar ## Comp $25,000 to $32,000/yr (part-time, 20 to 25 hrs/wk). Can scale to full-time for the right person. ## About the company HeartStamp is the greeting card and invitation platform for the AI era. Customers create hyper-personalized printed and Digital 3D cards with Stampy, our generative AI chat assistant, and what comes out is ready to send. We launched our US MVP this spring and have been shipping in weekly sprints ever since with a lean, distributed team. Quality is not bolted on at the end. It is how we build. ## How to apply Send a short note telling us one flaky test you have actually stabilized and one AI output problem an eval caught before users did. Skip the cover letter. Specific beats polished. If you have a Playwright repo or an eval rubric we can look at, send the link.
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗