CareerPlanGet AI match score →

Full Stack Engineer - Voice AI

🌐 Remote💼 Full-time🗓 2026-06-25

Core

Building server-side infrastructure for real-time, low-latency conversational AI voice agents and voicebots.

Role type

Senior Full Stack Engineer (Voice AI)

Builds

Production-grade voice agents, real-time audio pipelines, and LLM orchestration systems.

Domain

Real-time systems, Conversational AI, Voice Interfaces

Deliverable

production ML models | product features

Required skills

Backend engineering, Voice AI/voicebot system development, Real-time audio streaming, LLM orchestration, Prompt engineering, Python, Cloud infrastructure (AWS/GCP/Azure), CI/CD, Observability

Preferred skills

LLM frameworks (LangChain, LlamaIndex), AI agent frameworks (LiveKit Agents SDK, Pipecat), Telephony APIs (Twilio, Plivo), ASR/TTS integration, Open-source contributions

Technologies

LiveKit, WebRTC, Python, Node.js, Docker, Kubernetes, AWS, GCP, PostgreSQL, Redis, OpenAI, Anthropic, Deepgram, Whisper, ElevenLabs

Responsibilities

Design and iterate on voice AI agents for real-time conversations; Integrate and orchestrate LLMs as reasoning backbone; Manage real-time audio/media streams; Write and refine prompts for agent behavior; Build backend services connecting telephony/audio infra to business logic; Maintain cloud infrastructure and observability; Instrument systems for latency and SLAs.

Seniority

Mid-Senior, hands-on IC

Rewrite
## About the Role We are building the next generation of AI-powered voice experiences and are looking for a Full Stack Engineer who lives and breathes real-time, conversational AI. This is not a generic backend role, we need someone who has actually shipped voice agents in production: wrestled with latency, tuned prompts under load, and debugged the weird edge cases that only appear when a real user is talking to your system. You will own the server-side infrastructure that makes our voicebots fast, reliable, and smart , from LLM orchestration and real-time audio pipelines to deployment, scaling, and everything in between. ## What You'll Do - Design, build, and iterate on voice AI agents that handle real-time, low-latency conversations at scale. - Integrate and orchestrate LLMs (OpenAI,Gemini, Anthropic, open-source) as the reasoning backbone of our voice systems. - Work with real-time audio and communication frameworks — LiveKit, WebRTC, or equivalent — to manage media streams reliably. - Write, refine, and version-control prompts; apply prompt engineering techniques (chain-of-thought, few-shot, structured output) to improve agent behaviour. - Build robust backend services in Python that serve as the glue between telephony/audio infra, LLM APIs, and business logic. - Set up and maintain cloud infrastructure (AWS/GCP/Azure), CI/CD pipelines, and observability tooling — comfortable with Docker, Kubernetes, or similar. - Instrument systems for latency, uptime, and conversation quality; define and hit SLAs for real-time workloads. - Collaborate closely with product and ML teammates to translate requirements into production-grade voice features. ## What We're Looking For ### Non-Negotiable - 2 – 4 years of backend engineering experience, with meaningful time spent building voice AI or voicebot systems. - Demonstrable work on voice agents, could be LiveKit-based pipelines, Twilio integrations, VAPI, Retell AI, or a custom WebRTC stack. - Solid understanding of real-time audio streaming, concurrency challenges, and latency constraints unique to voice interfaces. ### Strong Plus - Experience with LLM frameworks, LangChain, LlamaIndex, or custom orchestration — and multi-step agent architectures. - Hands-on prompt engineering: system prompt design, function/tool calling, context management, and output validation. - Familiarity with AI agent frameworks such as LiveKit Agents SDK, Pipecat, Vocode, or similar. - Backend proficiency in Python and/or Node.js with production-grade API design (REST, WebSocket, gRPC). - DevOps literacy: Docker, CI/CD, cloud deployments, monitoring (Datadog, Grafana, or similar). - Experience designing systems that scale horizontally under bursty, real-time traffic. ## Nice to Have - Knowledge of ASR (Whisper, Deepgram, AssemblyAI) and TTS (ElevenLabs, Cartesia, Azure TTS) integration patterns. - Experience with telephony APIs (Twilio, Plivo, Vonage) or SIP/PSTN infrastructure. - Contributions to open-source AI tooling or published demos/side projects involving voice AI. ## Tech Stack We Work With - LiveKit Python, OpenAI / Anthropic APIs, LangChain, WebRTC, Docker, Kubernetes, AWS / GCP - PostgreSQL, Redis, Deepgram, Whisper, ElevenLabs / TTS ## Why Join Us - Work on genuinely hard, flexible work environment, cutting-edge problems at the intersection of real-time systems and AI. - Small, high-ownership team, your work ships to users, not to a backlog. - Competitive, market-benchmarked compensation with equity upside. - Flexible working environment; async-first culture with strong documentation habits. - Learning budget, conference access, and an opinionated internal knowledge base to help you grow fast.
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗