CareerPlanGet AI match score →

Qa Engineer

🌐 Remote💼 Full-time🗓 2026-07-24

Core

Ensure reliability of a platform spanning 368 apps, 64,000+ actions, SDKs, and packages used by developers, AI agents, and LLMs across major models and frameworks.

Role type

Senior SDET / Reliability Engineer (AI/LLM platform)

Builds

Automated test suites, load testing infrastructure, and monitoring pipelines for a developer platform and AI tooling.

Domain

Developer tools, AI/LLM integration, API reliability

Deliverable

production ML models | product features | dashboards & analysis

Required skills

End-to-end testing, REST API testing, Load/performance testing, JavaScript or Python, Test frameworks (Jest/Playwright/Pytest), AI agent harness usage, OAuth debugging, Schema validation, Rate limit handling, Failure triage

Preferred skills

Experience at platform/integration companies, Synthetic transaction testing, Model-agnostic test design, Open source contributions

Technologies

k6, Locust, Artillery, Claude Code, GPT, Claude, Gemini, HubSpot, Stripe, Shopify, Twilio

Responsibilities

Own end-to-end test coverage for 368+ platforms, Write and maintain tests for auth flows and action execution, Run tests against different LLMs and AI models, Build and run load tests, Build automated test suites for live APIs, Triage failures related to schema changes or auth issues, Monitor production action health and set up alerts

Seniority

Senior, hands-on IC

Rewrite
## About the role One's platform spans 368 apps, 64,000+ actions, a dashboard, SDKs, and additional packages — all of it used by developers, AI agents, and LLMs across every major model and framework. Each action is a live API call. APIs change without warning. Auth breaks. Schemas shift. Models behave differently. Rate limits bite. Your job is to make sure none of that reaches our users. ## Responsibilities - Own end-to-end test coverage for actions across 368+ platforms — from HubSpot and Stripe to Shopify, Twilio, and everything between. - Write and maintain end-to-end tests for our core system — auth flows, action execution, tool resolution, and the full request lifecycle. - Test our dashboard app and every additional package and product we own — end to end, across environments. - Run tests against different LLMs and AI models (GPT, Claude, Gemini, open-source models, and others) to ensure One works reliably regardless of what's calling it. - Test integrations with agent frameworks and AI tooling that uses One as its action layer. - Build and run load tests to validate system behavior under real traffic conditions and catch performance regressions before they hit production. - Build and maintain automated test suites that run against live APIs and catch regressions before they ship. - Triage failures fast — identify whether it's a schema change, auth issue, rate limit, model behavior, or a bug on our side. - Work directly with the integrations and platform teams to write tests as features are built, not after. - Monitor production action health and set up alerts that fire before customers notice. - Document failure patterns, maintain runbooks, and make on-call straightforward for the next person. ## Requirements - 3+ years in QA, SDET, or reliability engineering. - Deep technical understanding of how web systems work — HTTP, auth, APIs, distributed systems, and failure modes. - Strong experience writing automated tests against REST APIs — not just UI testing. - Experience testing LLM-powered products or AI tooling — you understand how model outputs and behavior add variability to test design. - Hands-on experience with load and performance testing tools (k6, Locust, Artillery, or similar). - The ability to write end-to-end tests that cover full system flows, not just isolated units. - Comfort with JavaScript or Python for test tooling and scripting. - Experience with test frameworks (Jest, Playwright, Pytest, or similar). - Hands-on experience with Claude Code, Codex, or a similar AI agent harness. This is not optional. - Clear written communication in English — we work async across time zones. - An instinct for what can go wrong, not just what should go right. ## Nice to have - You've worked at a platform or integration company where API reliability is core. - You've built monitoring pipelines or synthetic transaction tests in production. - You understand OAuth flows well enough to debug a 401 at 2am. - You've tested across multiple LLM providers and know how to design tests that are model-agnostic. - You've contributed to open source test tooling. ## How we work We use Claude Code and similar agent harnesses every day. If you already work this way, you will move 5x faster here. If you don't yet, you will learn quickly. Comfort building with AI is not optional. ## Why One - Remote, global, async by default. - Small team, real ownership, work that ships. - A category being built in real time. - Flexible setup and direct comms, always.
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗