CareerPlanGet AI match score →
💼 Full-time🗓 2026-06-25

Core

Co-designing and hardening a real-time voice infrastructure platform for multi-lingual translation, voice cloning, and contact center solutions.

Role type

Founding Engineer (Backend & Platform)

Builds

A voice infrastructure platform serving contact centers, AI agents, and public developers via SDK.

Domain

Voice infrastructure, real-time translation, audio streaming, and cloud security.

Deliverable

production ML models | product features | infrastructure

Required skills

Python (production-level), Async programming (FastAPI, asyncio, WebSockets), React with TypeScript, PostgreSQL (Row-Level Security), API design for public SDKs, CI/CD pipelines.

Preferred skills

Experience with WebRTC, WebSocket audio framing, jitter handling, streaming pipeline optimization, self-hosting LLM inference on GPU infrastructure.

Technologies

Python 3.12, FastAPI, WebSockets, asyncio, React, Vite, TypeScript, Tailwind CSS, PostgreSQL, GitHub Actions, Terraform, Pulumi.

Responsibilities

Design and harden the public API contracts and SDK, own the voice pipeline (STT, translation, TTS, voice cloning), manage multi-tenant security and compliance (C2PA, HIPAA, SOC 2), plan infrastructure migration to AWS, and lead the roadmap for self-hosting models.

Seniority

Founding, hands-on IC with path to CTO

Rewrite
## About the Role Founding Engineer — TransVoix (Voice Infrastructure Platform) TransVoix is a voice infrastructure platform for real-time multi-lingual translation maintaining a speaker's vocal identity with three product surfaces sharing one backend: human-to-human translation for contact centers, AI agent voice surfaces for voice AI platforms, and a public developer SDK with usage-based pricing. The platform is built around C2PA cryptographic voice provenance, BIPA-grade consent trails, HIPAA BAA readiness, and is pursuing a SOC 2 Type I program in 2026 with SOC 2 Type II as a 2027 roadmap goal. It is not a concept. It is a live product in closed beta with 40+ testers, an enterprise pilot in pre-launch with a Y Combinator–backed customer, bootstrapped to date, and patentable architecture currently in active conversation with IP counsel. I am looking for the founding engineer who can co-design and harden the API contracts external developers will hold us to for years. ## The Shift TransVoix originally launched as a real-time bidirectional voice translation app. As of May 2026, it has materially pivoted to a voice infrastructure platform. The pivot raises the stakes for the founding engineer materially. The work is no longer about shipping features in a translation app. It is about co-designing and hardening a platform whose API contracts paying customers will depend on years from now, while owning platform engineering through our Series A milestone. ## The Product Check out transvoix.ai and listen to the voice demo yourself. Happy to share a Beta Access code to demo the web application. Here is what you would be working on: - 6 languages (English, Spanish, French, Mandarin, Arabic, Hindi) with all 30 bidirectional pairs, not English-hub-only - Voice cloning where the listener hears your voice speaking their language with tone, prosody, and accent preserved - 2-party translated calls and group calls with colored avatars and real-time chat bubbles - Sub-second latency, 2,000+ automated tests across 100+ test files, multiple independent security pressure-test rounds, 4.72/5 translation quality - Enterprise pilot in pre-launch with a Y Combinator–backed customer ## The Stack - Backend: Python 3.12, FastAPI, WebSockets, asyncio — cloud-hosted, US region today with multi-region on the roadmap - Frontend: React, Vite, TypeScript, Tailwind CSS - Audio pipeline: streaming speech-to-text → LLM-based translation → neural text-to-speech with in-call voice cloning, all API-served today (provider and model specifics discussed on the screening call, not published) - Database and Auth: PostgreSQL with Row-Level Security hardened, atomic RPC patterns, ES256 and JWKS for JWT - Voice infrastructure: an AudioIngressAdapter Protocol abstracting ingress source from the pipeline contract — a single contract serves browser WebSocket, contact-center platforms, and future SIP/telephony ingress sources - Model serving (current): all inference is API-served via third-party providers. Self-hosting selected models is on the roadmap based on latency, cost, and fine-tuning requirements. - CI/CD: GitHub Actions on every PR, Dependabot, 2,000+ tests across 100+ files - Observability: error monitoring with PII-clean tagging discipline, structured logging, AST-pinned invariants enforcing security-critical patterns, product analytics ## What You Would Own - The voice pipeline: streaming STT → translation → TTS with sub-second latency, jitter buffers, AudioContext lifecycle, replaceTrack semantics, and the AudioIngressAdapter Protocol - The public SDK: whatever you ship here becomes a versioned contract external developers will hold us to. API design judgment is the gating skill, not coding speed. - The trust envelope: C2PA cryptographic voice provenance, BIPA-grade immutable consent rows with forbidden-key invariants, HIPAA BAA readiness, and the path from a 2026 SOC 2 Type I program toward SOC 2 Type II as a 2027 goal - The orchestration substrate: multi-tenant security where Row-Level Security and tenant isolation are load-bearing, signed audit logs, AST-pinned invariants for security-critical patterns - The infrastructure migration: we are on a managed cloud platform today and will migrate to AWS as we scale pilots. You will own the migration plan with Terraform or Pulumi for Infrastructure as Code, multi-region deployment, observability at scale, cost discipline, and zero-downtime cutover from the live platform serving real traffic. - The self-hosted model roadmap: today we run on third-party inference APIs. As we scale, latency, cost, and fine-tuning needs will push us toward self-hosting selected models. You will own the build-vs-buy analysis, the GPU infrastructure decisions, and the migration from API-served to self-hosted inference where it makes sense. - Engineering judgment under founder pressure: I will push for features, vendor swaps, and pilot accommodations. The right founding engineer pushes back with API versioning cost, regression test requirements, and a clear no when something would lock us into a contract we cannot break in six months. - Path to CTO based on fit ## Must-Have Skills - Python production-level depth. The entire backend is Python. - Async programming with FastAPI, asyncio, WebSockets. The audio pipeline is async-heavy. - React with TypeScript. The frontend is React/Vite/TypeScript. You need to be comfortable here. - PostgreSQL with Row-Level Security, migrations, query optimization. - API integration with multiple external providers. - CI/CD: GitHub Actions, deployment pipelines, testing infrastructure. ## What We Need You to Have Done Before - Production voice or audio pipeline work where latency was load-bearing. Not theoretical familiarity. Specifically WebRTC, WebSocket audio framing, jitter handling, streaming pipeline optimization. If you have shipped production audio infrastructure on a contact-center, telephony, or real-time-media platform, that is the depth we mean. - Public SDK or API design where external developers depended on your decisions. Either you have shipped a public SDK or designed an API that external developers relied on.
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗