CareerPlanGet AI match score →

Audio Inference Engineer, Model Efficiency

New York🌐 Remote💼 Full-time🗓 2025-11-07 → 2026-07-31

Core

Build reliable machine learning systems and optimize audio inference serving efficiency using innovative techniques.

Role type

Senior IC machine-learning engineer (audio inference efficiency)

Builds

Real-time and streaming audio inference systems

Domain

AI / Audio Processing / Distributed Systems

Deliverable

production ML models

Required skills

C++, Python, deep learning models for audio/speech, high-performance inference system development, system bottleneck identification, creative solution design for audio processing

Preferred skills

GPU programming, low-level system optimization, model parallelization over multiple GPUs, duplex real-time streaming architectures, machine learning framework internals (PyTorch, TensorFlow), inference frameworks (vLLM, SGLang, Tensort-LLM), sequence modeling (transformers for audio/speech), end-to-end audio pipeline optimization

Technologies

C++, Python, PyTorch, TensorFlow, vLLM, SGLang, Tensort-LLM

Responsibilities

Advance core audio model serving metrics (latency, throughput, quality), dive deep into systems to identify bottlenecks, deliver creative solutions for audio processing and streaming workloads, collaborate with training and serving infrastructure teams for seamless integration

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗