Inference Engineer
Core
Building and optimising realtime TTS streaming infrastructure and runtime systems for low-latency conversational speech models to power enterprise voice experiences.
Role type
Senior IC machine-learning engineer (realtime inference systems)
Builds
Production-grade realtime speech inference systems serving hundreds of millions of conversations
Domain
Voice AI / Realtime speech infrastructure
Deliverable
production ML models
Required skills
Realtime inference optimisation, scheduler design, GPU utilisation, concurrency optimisation, dynamic batching, TensorRT, Triton, ONNX Runtime, CUDA, Rust, C++, Python, Kubernetes, vLLM, CUDA Graphs
Preferred skills
Experience with heterogeneous GPU environments (NVIDIA/AMD), speculative decoding, KV cache management, kernel-level profiling
Technologies
TensorRT, Triton, ONNX Runtime, vLLM, CUDA, Kubernetes, AWS, Rust, C++, Python
Responsibilities
Building and optimising realtime TTS streaming infrastructure, Improving scheduler and batching systems for production workloads, Reducing TTFA/TTFB while maintaining speech quality and stability, GPU profiling and identifying kernel-level bottlenecks, Optimising TensorRT, Triton, ONNX Runtime, and custom serving systems, Managing KV cache systems, speculative decoding, and streaming inference
Seniority
Senior, hands-on IC