CareerPlanSign in

LLM Inference

United States, Multiple Locations, Multiple Locations💼 Full-time🗓 2025-11-27 → 2026-09-26

Core

Implement frontier AI research ideas to improve LLM inference performance, debug bottlenecks, and build tools for distributed systems.

Role type

Senior IC LLM inference engineer

Builds

High-performance inference systems and debugging tools for generative AI models

Domain

Generative AI / Large Language Models / Distributed Systems

Deliverable

production ML models

Required skills

C/C++/C#/Java/JavaScript/Python, generative AI, distributed computing, Python ecosystem (uv, pybind/nanobind, FastAPI), large scale production inference, GPU kernel programming, PyTorch benchmarking and optimization, vLLM, SGLang, JAX scaling

Technologies

PyTorch, vLLM, SGLang, JAX, uv, pybind, nanobind, FastAPI

Responsibilities

Implement frontier AI research ideas, introduce new systems and tools to improve inference performance, build tools to debug performance bottlenecks and numeric instabilities, establish processes to enhance team productivity, deliver work iteratively to users

Seniority

Senior, hands-on IC

Sourced via microsoft · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.