LLM Inference
Core
Implement frontier AI research ideas to improve LLM inference performance, debug bottlenecks, and build tools for distributed systems.
Role type
Senior IC LLM inference engineer
Builds
High-performance inference systems and debugging tools for generative AI models
Domain
Generative AI / Large Language Models / Distributed Systems
Deliverable
production ML models
Required skills
C/C++/C#/Java/JavaScript/Python, generative AI, distributed computing, Python ecosystem (uv, pybind/nanobind, FastAPI), large scale production inference, GPU kernel programming, PyTorch benchmarking and optimization, vLLM, SGLang, JAX scaling
Technologies
PyTorch, vLLM, SGLang, JAX, uv, pybind, nanobind, FastAPI
Responsibilities
Implement frontier AI research ideas, introduce new systems and tools to improve inference performance, build tools to debug performance bottlenecks and numeric instabilities, establish processes to enhance team productivity, deliver work iteratively to users
Seniority
Senior, hands-on IC