CareerPlanGet AI match score →

Member of Technical Staff - Inference Research

New York💼 Full-time🗓 2026-07-05 → 2026-08-01

Core

Building a platform to serve LLMs at scale, focusing on optimizing inference cost and latency through research into serving stack primitives.

Role type

Inference Research Engineer (Systems/LLM)

Builds

Production-ready inference primitives, disaggregated prefill/decode, quantization, KV-cache management, and autoscaling for serverless traffic.

Domain

Cloud Infrastructure / Large Language Model Serving

Deliverable

production ML models

Required skills

LLM inference research, systems programming, kernel development, quantization (FP8, INT4), KV-cache management, autoscaling, speculative decoding, distributed systems

Preferred skills

Experience with ZLab DFlash, SGLang, Flash Attention 4, open-source contributions

Technologies

ZLab, SGLang, Flash Attention 4, FP8, INT4

Responsibilities

Own end-to-end inference research bets, train custom speculators against production traffic, collaborate with customers to deploy and tune models, expand collaborations with outside research labs, work with engineering to turn research into products, help shape the research agenda

Seniority

Mid-Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗