Member of Technical Staff, Inference
Core
Build and optimize the vLLM inference engine to serve larger and more complex LLM and diffusion models across diverse hardware.
Role type
Senior IC inference runtime engineer
Builds
High-performance inference runtime for LLMs and diffusion models
Domain
AI Infrastructure / Large Language Models
Deliverable
production ML models
Required skills
Python, PyTorch internals, Transformer architectures, LLM inference systems, Research paper implementation, Complex codebase debugging
Preferred skills
KV-cache memory management, Prefix caching, Hybrid model serving, RL frameworks, Multimodal inference, Open-source contributions
Technologies
vLLM, TensorRT-LLM, SGLang, TGI, PyTorch
Responsibilities
Optimize model execution across diverse hardware and architectures, Implement model architectures and inference techniques from research papers, Contribute performant and maintainable code to the vLLM codebase
Seniority
Senior, hands-on IC