CareerPlanSign in

Software Engineer- Inference Platform

San Francisco💼 Full-time🗓 2026-10-05 → 2026-10-07

Core

Build the distributed runtime and orchestration systems that power large-scale LLM inference, ensuring models are fast, reliable, and cost-efficient for customers.

Role type

Senior IC distributed systems engineer (inference platform)

Builds

Distributed runtime for LLM inference, Model APIs, and multi-cloud capacity management

Deliverable

production ML models

Required skills

distributed systems, backend infrastructure, large-scale APIs, low-latency services, infrastructure profiling, SLO management, Kubernetes, observability, release automation

Preferred skills

LLM inference engines (vLLM, SGLang, TensorRT-LLM, TGI, Dynamo), GPU workloads, service meshes, API gateways, distributed scheduling, open-source contributions

Technologies

Kubernetes, vLLM, SGLang, TensorRT-LLM, TGI, Dynamo

Responsibilities

Build infrastructure and orchestration systems for deploying and running distributed LLM inference; Design and operate Model APIs with advanced capabilities like structured outputs and tool calling; Implement platform fundamentals including API versioning, metering, and authentication; Instrument deep observability and build benchmarks for speed and reliability; Debug and harden production systems spanning networking and GPU workloads; Partner with performance teams to optimize inference; Own projects end-to-end from architecture to deployment.

Seniority

Senior, hands-on IC