Founding AI Engineer
About the Role
This is a founding-level AI engineering role at an early-stage B2B SaaS pricing intelligence startup based in San Francisco. You'll be one of the first engineers on a small, product-focused team — joining at a formative moment to shape the technical direction of an AI-powered pricing platform used in high-stakes B2B deals.
Your work sits at the intersection of LLM infrastructure, evaluation systems, and revenue-critical product outcomes. You'll build the feedback loops, eval harnesses, and agent tooling that allow a pricing AI to earn trust, remain auditable, and drive measurable business impact for customers.
What You'll Do
• Design and build eval harnesses and benchmarks that use tracked pricing outcomes as ground truth, supporting release gating against a typed ontology of pricing entities.
• Systematize and automate expert review workflows currently performed manually.
• Develop AI personas that simulate B2B buying committees and behavioral effects, leveraging usage data and call transcripts.
• Automate persona training pipelines that are today manual processes.
• Own LLM routing across providers (e.g., Anthropic, Google) with explicit cost, latency, and quality tradeoffs.
• Maintain infrastructure and data residency boundaries (e.g., ensuring EU model calls remain within the EU; SOC2/GDPR compliance).
• Extend the MCP server used by LLM agents — including customer-facing agents — so that features are agent-driven, not just UI-rendered.
• Work within a typed ontology of pricing entities (e.g., Pydantic models for SKU, Proposition, Persona, Quote) so model outputs are structured and auditable.
• Identify and remediate systemic latency, data drift, and cold-start issues in the pricing loop.
What We're Looking For
Dealbreakers
• 8+ years of engineering experience with strong, recent production LLM depth — including shipping and owning LLM-powered features end-to-end after launch.
• Must be authorized to work in the United States; visa sponsorship is not available.
• Hands-on experience building evals and observability for LLM systems — creating eval harnesses, baselining prompts, and supporting release gating.
Required
• Experience owning LLM infrastructure, including routing across multiple model providers with cost, latency, and quality tradeoffs.
• Proficiency with structured data models and typed ontologies (e.g., Pydantic) to ensure model outputs are structured and auditable.
• Full end-to-end ownership of LLM-powered features — from development through production monitoring.
• Strong communication skills to explain non-deterministic systems to non-technical clients and stakeholders.
• Product-engineer instincts: ability to scope and deliver pragmatic solutions in a fast-moving environment.
Nice to Have
• Experience with MCP or building tools/integrations for LLM agents.
• Familiarity with platforms such as LangChain, LlamaIndex, Braintrust, or OpenRouter.
• Experience in domains where pricing, billing, or payments accuracy and auditability are required.
• Familiarity with data residency or compliance controls (SOC2, GDPR).
Compensation & Benefits
• Salary: $225,000 – $255,000 USD annually
• Founding-team equity opportunity at an early-stage, high-growth startup
Location
This role is on-site in San Francisco, CA. Candidates must be based in or willing to relocate to the San Francisco Bay Area.