Distributed LLM Inference Engineer
Core
Building high-performance distributed systems for large-scale LLM inference, optimizing Ray and integrating with open-source engines like vLLM to serve open-source users and enterprise customers.
Role type
Senior IC distributed systems engineer (LLM inference)
Builds
Scalable batch and online inference solutions for Ray ecosystem users and Anyscale customers
Domain
AI infrastructure, distributed systems, large language model serving
Deliverable
production ML models
Required skills
distributed systems, deep learning frameworks (PyTorch), ML inference optimization, open-source software integration, state-of-the-art research implementation
Preferred skills
ML systems knowledge, Ray experience, LLM engine expertise (vLLM, TensorRT-LLM), deep learning framework contributions (PyTorch, TensorFlow), deep learning compiler contributions (Triton, TVM, MLIR), GPU/CUDA experience
Technologies
Ray, vLLM, TensorRT-LLM, PyTorch, TensorFlow, Triton, TVM, MLIR, CUDA
Responsibilities
Iterate with product teams to ship end-to-end batch and online inference solutions at high scale; integrate Ray Data and LLM engines to achieve low-cost large-scale ML inference; integrate with and contribute to open-source software like vLLM; implement and extend best practices from the research community
Seniority
Senior, hands-on IC