Senior Inference Engineer, AGI
Core
Senior Inference Engineer owning the full-stack inference path for real-time multimodal conversational AI, co-designing model architectures for servability and building low-latency runtimes.
Role type
Senior Inference Engineer (Real-time Multimodal AI)
Builds
Real-time streaming serving runtime, offline training/RL infrastructure, and high-performance inference kernels.
Domain
AI/ML, Speech & Audio, Real-time Systems
Deliverable
production ML models
Required skills
Neural deep learning methods, Transformer architectures, GPU performance optimization, Inference optimization (quantization, speculative decoding, operator fusion), High-performance kernel development, Real-time streaming system design, Multi-GPU inference, Distributed training pipelines
Preferred skills
Custom GPU kernel authoring (CUTLASS, Triton, CUDA), Model compression techniques, Reinforcement learning infrastructure, Distributed training (NCCL, NVLink), Speech-to-speech/audio generative models, Open-source inference contributions
Technologies
vLLM, PyTorch, TensorRT-LLM, CUDA, CUTLASS, Triton, Nsight Compute, FlashAttention, NCCL, NVLink
Responsibilities
Partner with scientists to co-design inference-friendly model architectures; Implement and optimize inference paths for large-scale multimodal models; Develop high-performance kernels for critical operations; Own the real-time serving path for streaming multimodal conversational AI; Build offline inference systems for post-training and reinforcement learning; Profile end-to-end performance to eliminate bottlenecks.
Seniority
Senior, hands-on IC