Rewrite
## About the role
We're looking for an AI Engineer to own and improve the core AI systems that power our product.
This is not a basic prompt engineering role. You'll work on production AI systems that run live voice conversations, manage agent memory, evaluate conversation quality, and improve agent performance at scale.
You'll work across Python backend systems, LLM orchestration, voice AI pipelines, agent workflows, memory systems, and AI evaluation infrastructure.
## What You'll Work On
### Real-time Voice AI
- Build and optimize low-latency streaming speech pipelines
- Work with speech-to-text, LLM, and text-to-speech pipelines
- Improve turn detection, interruption handling, and response timing
- Track and optimize latency across speech, reasoning, and synthesis stages
- Work with voice AI tools and APIs such as Deepgram, ElevenLabs, OpenAI, Gemini, or similar platforms
### Agentic AI Systems
- Build multi-step AI agents with tool calling and workflow execution
- Design dynamic instructions, context injection, and guardrails
- Improve agent reliability in live conversations
- Build agents that can reason, call tools, update data, extract information, and take conditional actions
### LLM Orchestration
- Work with LLM APIs such as OpenAI, Gemini, Anthropic, or similar providers
- Build routing, fallback, retry, and error-handling logic for LLM calls
- Optimize prompts, system instructions, context windows, and structured outputs
- Improve reliability, latency, and cost-efficiency of LLM-powered workflows
### Memory & Context Engineering
- Build memory systems using call history, contact profiles, and retrieval-based context
- Improve context retrieval pipelines for accurate and low-latency responses
- Work with embeddings, vector search, semantic retrieval, and RAG pipelines
- Extract useful facts, preferences, and follow-up signals from transcripts
### AI Evaluation & Quality
- Build evaluation systems for scoring agent performance
- Design rubrics for goal completion, objection handling, accuracy, script adherence, and conversation quality
- Build LLM-as-judge and automated quality-checking systems
- Detect quality regressions after model, prompt, or workflow changes
### Post-call Intelligence
- Build automated summarization, sentiment analysis, fact extraction, and outcome detection
- Extract patterns from successful and failed conversations
- Turn call data into actionable insights for users
- Build structured data extraction pipelines from call transcripts
### Backend & Infrastructure
- Build and maintain Python backend services
- Work with async APIs, background jobs, task queues, webhooks, and cloud-native deployments
- Improve reliability, observability, latency, and cost efficiency of AI workloads
- Work with Docker, Kubernetes, databases, queues, and production monitoring systems
## You're a Fit If You
- Have strong Python experience
- Have shipped LLM-powered systems in production
- Have worked with OpenAI, Gemini, Anthropic, or similar LLM APIs
- Have worked with voice AI APIs such as Deepgram, ElevenLabs, or similar STT/TTS platforms
- Have built multi-step agents with tool calling
- Understand real-time constraints like latency, streaming, and interruption handling
- Have experience with RAG, vector search, embeddings, and context engineering
- Know how to evaluate AI quality beyond manual checks
- Are comfortable with async Python, APIs, queues, background workers, and distributed systems
- Understand prompt engineering, structured outputs, function calling, and context window management
- Like owning systems end to end in a fast-moving environment
## Required Skills
- Python
- LLM-powered application development
- OpenAI / Gemini / Anthropic or similar LLM APIs
- Deepgram / ElevenLabs or similar voice AI APIs
- AI agents and tool calling
- Prompt engineering
- Context engineering
- RAG pipelines
- Vector search and embeddings
- Structured output extraction
- Async backend systems
- APIs, webhooks, queues, and background jobs
- Real-time streaming systems
- LLM evaluation and AI quality monitoring
- Docker, Kubernetes, and cloud infrastructure
- Production debugging, logging, and observability
## Bonus Points
- Experience with real-time voice AI or low-latency audio systems
- Experience building AI evaluation frameworks or LLM observability tools
- Experience with speech-to-text and text-to-speech pipelines
- Experience with multi-provider LLM routing and fallback systems
- Experience with preference optimization, feedback loops, or strategy extraction
- Experience with React/TypeScript for contributing to product dashboards
- Early-stage startup experience
## What We Offer
- Direct collaboration with technical leadership
- High autonomy and end-to-end ownership
- Opportunity to work on production AI systems used by real customers
- Work at the frontier of real-time voice AI and agent systems
- Competitive compensation based on experience
## To Apply
- Send your resume plus a short note on one AI system you've built and what made it technically hard.
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.