CareerPlanGet AI match score →

Member of Technical Staff — Audio and Voice AI

In office • New York City+1💼 Full-time🗓 2026-06-24

Core

Design, build, and deploy AI-powered voice and audio systems to solve real-world financial operations challenges, including intelligent voice agents and speech-driven workflows.

Role type

Senior IC machine-learning engineer (audio and voice AI)

Builds

Production-grade, real-time voice experiences and audio-driven workflow automation for enterprise finance teams.

Domain

Financial services (accounts receivable, collections) + Audio AI / Speech AI

Deliverable

production ML models | product features

Required skills

software engineering, applied AI/ML, speech systems, audio systems, LLM fine-tuning, real-time streaming inference, LLMOps, pipeline orchestration, system monitoring

Preferred skills

conversational AI, multimodal AI, financial domain knowledge

Technologies

speech-to-text, text-to-speech, conversational AI platforms, LLMs

Responsibilities

Design and ship production-ready audio and voice-based AI features; fine-tune and optimize speech, audio, and LLM-based models; build end-to-end AI/ML pipelines for audio ingestion and streaming inference; establish evaluation and monitoring frameworks for voice and text-based AI systems; develop AI-powered voice automations for financial workflows; partner with cross-functional teams to translate business needs into voice AI solutions.

Seniority

Senior, hands-on IC

Rewrite
## About the Role Stuut is transforming accounts receivable for B2B companies—making collections smarter and faster for companies that have historically relied on manual processes that are labor intensive and costly. Our platform is gaining traction with finance teams across industrials, chemicals, and manufacturing sectors from Fortune 10 brands to scaling midmarkets. We're backed by top-tier investors including a16z, Khosla, Activant, 1984 Ventures and Page One. We're hiring a Member of Technical Staff – Audio and Voice AI Systems to design, build, and deploy AI-powered voice and audio systems that solve real-world financial operations challenges. You'll take state-of-the-art research in speech, audio, and multimodal AI and translate it into production-grade, real-time voice experiences that deliver measurable customer impact. From intelligent voice agents that interact with customers to audio-driven workflow automation and speech-based data extraction, you'll create scalable, reliable AI systems that integrate seamlessly into Stuut's platform. Your work will directly shape how finance teams and their customers interact with Stuut through voice. This is a hands-on role for an engineer who thrives at the intersection of audio AI, real-time systems, and practical business impact—turning cutting-edge models into trusted, delightful voice experiences for enterprise finance workflows. ## Responsibilities - Build & Deploy Voice AI Systems: design and ship production-ready audio and voice-based AI features, including real-time voice agents and speech-driven workflows. - Craft High-Quality Voice UX: use modern speech-to-text, text-to-speech, and conversational AI platforms to create natural, responsive, and emotionally aware voice experiences tailored to financial use cases. - Adapt & Fine-Tune Audio and Multimodal Models: fine-tune and optimize speech, audio, and LLM-based models for accuracy, latency, and reliability in real-world environments. - Engineer Real-Time, Scalable AI Pipelines: build end-to-end AI/ML pipelines spanning audio ingestion, streaming inference, orchestration, and monitoring with enterprise-grade availability and performance. - Establish Evaluation & Monitoring Frameworks (LLMOps): design rigorous evaluation systems to measure quality, latency, accuracy, drift, and business outcomes for voice and text-based AI systems. - Automate Financial Workflows via Voice: develop AI-powered voice automations that reduce manual effort in collections, reconciliation, and customer communication. - Collaborate Cross-Functionally: partner with Product, Engineering, Design, and customers to translate business needs into effective, user-centered voice AI solutions. - Measure & Communicate Impact: define success metrics and continuously improve AI systems based on real-world usage and customer feedback. ## Requirements - Have 5+ years of software engineering experience, with 2+ years focused on applied AI/ML, speech, or audio systems in production
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗