CareerPlanGet AI match score →

Founding Harness Engineer

💼 Full-time💰 $130,000–$130,000🗓 2026-07-31

Core

Designing the intelligence layer between AI models and product, including system prompts, tool integrations, context management, and evaluation frameworks to enable reliable agent workflows.

Role type

Founding Harness Engineer (AI Agent Infrastructure)

Builds

Production-ready agent systems with measurable quality and reliability

Domain

Generative AI, LLM orchestration, content generation

Deliverable

production ML models

Required skills

System prompt engineering, LLM evaluation framework design, context window management, multi-step workflow architecture, A/B testing, technical writing

Preferred skills

Background in English, literature, creative writing, or humanities

Technologies

LLMs, agent frameworks, evaluation metrics

Responsibilities

Build and refine system prompts and tool integrations; Design evaluation frameworks (LLM-as-judge, deterministic tests); Architect abstraction layers between product and model capabilities; Own prompt versioning and experimentation; Collaborate with product/engineering to identify failure signals and improve agent quality

Seniority

Founding, hands-on IC with architectural ownership

Rewrite
## About the role Harness Engineer Location: San Francisco (in-person, in-office, full-time ONLY) ## About Virio Virio is solving an unsolved problem: how do you write perfect content? The best human strategist researches the space, finds the angle, and knows your goals and constraints — but they're capped by the hours in a day and one human career's worth of expertise. An agent isn't. Crafting "perfect" content has three components: - Research: to identify and collect the best inputs - Synthesis: to devise the strategy and execute on it - Verification: RL verification of content quality and efficacy We're backed by operators from LinkedIn, YouTube, HubSpot, Rippling, and Google. The team is almost 100% ex-founders and VCs from Yale, UC Berkeley, Stripe, and more. And we've consecutively added $1M ARR/mo because everyone and their co-founder wants what we're selling. This problem doesn't have a playbook. We're looking for a Harness Engineer to obsess over AI-for-writing and own that layer end to end: the data feedback loop, evals, and harness design that turn a subjective judgment into something we can measure and improve — so day by day, we can get ever closer to solving content. ## What You'll Do As a Harness Engineer, you'll own the intelligence layer that sits between our AI models and product—the system prompts, tool definitions, context management, and evaluation framework that makes agents actually work in production. Specifically: - Build and refine the system prompts, tool integrations, and context windows that shape how our models behave across our platform—you'll see how small prompt tweaks cascade into measurable product impact. (Nice to have: background in English, literature, creative writing, or humanities) - Design the evaluation framework (LLM-as-judge evals, deterministic tests, production monitoring) that lets us ship with confidence. You'll define what success looks like and build the metrics to measure it. - Architect the abstraction layer between our product (file systems, artifacts, skills) and the model's capabilities—making complex multi-step workflows feel natural to the model and reliable to users. - Own prompt versioning, experimentation, and iteration. You'll A/B test prompt variations, measure their impact on agent quality, and ship improvements across our production system. - Collaborate with product and engineering to identify signal—where our agents are failing, where users are stuck—and translate that into prompt and architecture improvements. ## What Success Looks Like After 30 days: You understand our agent stack—the system prompts, available tools, context window strategy, and current failure modes. You've shipped your first prompt improvement and measured its impact. After 90 days: You've designed and implemented our core evaluation framework. Our team uses it as the source of truth for agent quality. You've shipped 2-3 meaningful prompt iterations that improved measurable outputs, and you're comfortable making architecture trade-offs. After 6 months: You own the harness layer end-to-end. You've architected improvements that reduced hallucination, improved context retention, or increased model reliability. You're driving architectural decisions about how we expose capabilities to the product team. You're writing technical content about how we've solved these problems—attracting the next generation of harness engineers to the company. ## Traits That Thrive at Virio - High agency & self-direction – Acts without waiting for requirements or detailed instructions - Small-team intensity comfort – Comfortable working intensely and closely with a small, high-output team - Unorthodox problem solving – Sees non-obvious paths forward and challenges conventional approaches - Quality-driven execution – Detail-oriented with high standards for correctness and quality - Low ego & accountability – Takes full ownership and accountability for outcomes ## What we offer - Work directly with the founders - Competitive total compensation aligned to impact and ownership - Meaningful equity upside as part of the founding team - Medical, dental, and vision insurance - 401(k) plan - All working meals covered - Relocation support for San Francisco - Company-wide annual off-site - High ownership, high trust, and fast career progression alongside world-class client partners - In-person culture with deep commitment to excellence This is a full-time, on-site role based in San Francisco. We do not offer hybrid or remote positions.
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗