ML Runtime Engineer (Mid-Level and Senior)
Core
Building AI acceleration hardware and runtime stacks to run large language models 100x faster.
Role type
Senior ML Runtime Engineer (IC)
Builds
Scalable inference engines and runtime stacks for AI accelerators
Domain
AI Infrastructure / Hardware-Software Integration
Deliverable
production ML models
Required skills
ML inference at scale, multi-user serving, paged attention, inference engines (vLLM, SGLang), transformer architecture internals, software engineering
Preferred skills
Rust, building inference engines from scratch
Technologies
vLLM, SGLang, Rust
Responsibilities
Integrate AI acceleration hardware with inference engines, research and implement KV cache management technologies, design and build scalable reference inference engines, focus on transformer ML architecture internals, shape the direction of the runtime stack
Seniority
Senior, hands-on IC