Software Engineer, Model Runtime
Core
Design and implement a production-grade LLM inference runtime for frontier models running on custom OpenAI silicon, optimizing throughput, latency, and hardware utilization.
Role type
Senior IC systems software engineer (LLM inference runtime)
Builds
High-performance inference engine components including scheduling, continuous batching, memory management, and distributed execution strategies.
Domain
AI hardware, distributed systems, compilers, and LLM inference
Deliverable
production ML models
Required skills
C++, Rust, Python, distributed systems, compilers, kernels, LLM inference (prefill/decode, batching, KV-cache, model parallelism), performance optimization, profiling, debugging
Preferred skills
Experience with model-serving infrastructure, hardware-software co-design, quantitative reasoning about compute/memory/communication, clean abstraction design
Technologies
C++, Rust, Python, custom silicon, vLLM, SGLang
Responsibilities
Design and implement LLM inference runtime; build scheduling, batching, memory, and KV-cache management; develop distributed execution strategies; optimize latency and throughput; partner with kernel/compiler/silicon teams; enable new model features; create profiling and observability tools; debug correctness and performance issues; translate workload insights to silicon requirements
Seniority
Senior, hands-on IC