CareerPlanGet AI match score →

Distributed LLM Inference Engineer

San Francisco💼 Full-time🗓 2026-05-27 → 2026-07-31

Core

Building high-performance distributed systems for large-scale LLM inference, optimizing Ray and integrating with open-source engines like vLLM to serve open-source users and enterprise customers.

Role type

Senior IC distributed systems engineer (LLM inference)

Builds

Scalable batch and online inference solutions for Ray ecosystem users and Anyscale customers

Domain

AI infrastructure, distributed systems, large language model serving

Deliverable

production ML models

Required skills

distributed systems, deep learning frameworks (PyTorch), ML inference optimization, open-source software integration, state-of-the-art research implementation

Preferred skills

ML systems knowledge, Ray experience, LLM engine expertise (vLLM, TensorRT-LLM), deep learning framework contributions (PyTorch, TensorFlow), deep learning compiler contributions (Triton, TVM, MLIR), GPU/CUDA experience

Technologies

Ray, vLLM, TensorRT-LLM, PyTorch, TensorFlow, Triton, TVM, MLIR, CUDA

Responsibilities

Iterate with product teams to ship end-to-end batch and online inference solutions at high scale; integrate Ray Data and LLM engines to achieve low-cost large-scale ML inference; integrate with and contribute to open-source software like vLLM; implement and extend best practices from the research community

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗