Research Engineer - Inference
Core
Deploy and optimize frontier AI models in production for real-time, streaming workloads, turning research breakthroughs into scalable products.
Role type
Senior IC research engineer (inference optimization)
Builds
High-performance serving systems and tooling for real-time AI audio models
Domain
AI/ML inference, real-time streaming systems
Deliverable
production ML models
Required skills
GPU programming, inference optimization, CUDA, Triton, TensorRT, vLLM, SGLang, profiling, bottleneck elimination, custom kernel development
Preferred skills
Experience with latency-sensitive applications, model quantization, distillation, KV-cache optimization, batching strategies
Technologies
CUDA, Triton, TensorRT, vLLM, SGLang
Responsibilities
Deploy state-of-the-art models to production; optimize inference performance across the stack; build and tune high-performance serving systems; create tooling for researchers to ship models quickly
Seniority
Senior, hands-on IC