CareerPlanSign in

Member of Technical Staff, Inference

San Francisco💼 Full-time💰 $200,000–$200,000🗓 2026-09-16 → 2026-09-28

Core

Build and optimize the vLLM inference engine to serve larger and more complex LLM and diffusion models across diverse hardware.

Role type

Senior IC inference runtime engineer

Builds

High-performance inference runtime for LLMs and diffusion models

Domain

AI Infrastructure / Large Language Models

Deliverable

production ML models

Required skills

Python, PyTorch internals, Transformer architectures, LLM inference systems, Research paper implementation, Complex codebase debugging

Preferred skills

KV-cache memory management, Prefix caching, Hybrid model serving, RL frameworks, Multimodal inference, Open-source contributions

Technologies

vLLM, TensorRT-LLM, SGLang, TGI, PyTorch

Responsibilities

Optimize model execution across diverse hardware and architectures, Implement model architectures and inference techniques from research papers, Contribute performant and maintainable code to the vLLM codebase

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 850,000+ jobs from 20+ sources.