Machine Learning Engineer, Ops
Core
Build and scale low-latency, reliable inference infrastructure for generative audio models (TTS, voice conversion, ASR) to bridge research and production.
Role type
Senior MLOps Engineer (Inference Infrastructure)
Builds
High-performance model serving systems for streaming and batch inference
Domain
Generative AI, Audio Models, Cloud Infrastructure
Deliverable
production ML models
Required skills
Kubernetes (K8S) orchestration, CI/CD automation, Python or Go, GPU-accelerated inference, performance profiling, model architecture knowledge (TTS/ASR)
Preferred skills
Triton Inference Server, vLLM-Omni
Responsibilities
Design and maintain inference infrastructure, implement high-performance inference engines, orchestrate service deployments with autoscaling, develop CI/CD pipelines, monitor production systems for latency and resource utilization, optimize inference performance for streaming and batch applications
Seniority
Senior, hands-on IC
