Member of Technical Staff - Inference Research
Core
Building a platform to serve LLMs at scale, focusing on optimizing inference cost and latency through research into serving stack primitives.
Role type
Inference Research Engineer (Systems/LLM)
Builds
Production-ready inference primitives, disaggregated prefill/decode, quantization, KV-cache management, and autoscaling for serverless traffic.
Domain
Cloud Infrastructure / Large Language Model Serving
Deliverable
production ML models
Required skills
LLM inference research, systems programming, kernel development, quantization (FP8, INT4), KV-cache management, autoscaling, speculative decoding, distributed systems
Preferred skills
Experience with ZLab DFlash, SGLang, Flash Attention 4, open-source contributions
Technologies
ZLab, SGLang, Flash Attention 4, FP8, INT4
Responsibilities
Own end-to-end inference research bets, train custom speculators against production traffic, collaborate with customers to deploy and tune models, expand collaborations with outside research labs, work with engineering to turn research into products, help shape the research agenda
Seniority
Mid-Senior, hands-on IC