Backend Engineer- Inference Services
Core
Design and implement secure, robust, and scalable services for speech processing, efficient distributed compute orchestration, and optimized scheduling for Deepgram's inference platform.
Role type
Senior Backend Software Engineer (Inference Services)
Builds
Real-time speech-to-text (STT), text-to-speech (TTS), and voice agent APIs
Domain
Voice AI / Speech Processing / Distributed Systems
Deliverable
production ML models
Required skills
Rust, C, C++, Python, UNIX-style systems, version control (git), networking, high performance computing, latency optimization, memory optimization
Preferred skills
modern machine learning frameworks (Torch), audio processing, CNNs, RNNS, transformers
Responsibilities
Improve core inference services including networking, speech processing, audio transcoding, and latency/memory optimization; Develop processes for measuring, building, and optimizing services; Debug complex system issues involving networking, scheduling, and HPC; Rapidly customize backend services for customer needs; Partner with Product to design and implement new services end to end