CareerPlanSign in

Software Engineer, Model Runtime

San Francisco💼 Full-time🗓 2026-08-24 → 2026-09-25

Core

Design and implement a production-grade LLM inference runtime for frontier models running on custom OpenAI silicon, optimizing throughput, latency, and hardware utilization.

Role type

Senior IC systems software engineer (LLM inference runtime)

Builds

High-performance inference engine components including scheduling, continuous batching, memory management, and distributed execution strategies.

Domain

AI hardware, distributed systems, compilers, and LLM inference

Deliverable

production ML models

Required skills

C++, Rust, Python, distributed systems, compilers, kernels, LLM inference (prefill/decode, batching, KV-cache, model parallelism), performance optimization, profiling, debugging

Preferred skills

Experience with model-serving infrastructure, hardware-software co-design, quantitative reasoning about compute/memory/communication, clean abstraction design

Technologies

C++, Rust, Python, custom silicon, vLLM, SGLang

Responsibilities

Design and implement LLM inference runtime; build scheduling, batching, memory, and KV-cache management; develop distributed execution strategies; optimize latency and throughput; partner with kernel/compiler/silicon teams; enable new model features; create profiling and observability tools; debug correctness and performance issues; translate workload insights to silicon requirements

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.