CareerPlanSign in

Software Engineer, Inference

San Francisco, CA💼 Full-time🗓 2026-10-05 → 2026-10-07

Core

Design and build large-scale inference systems for AI agents, managing routing, capacity, and optimization across self-hosted models and third-party providers.

Role type

Senior IC distributed systems engineer (AI inference)

Builds

High-performance inference serving stack, routing layers, and GPU infrastructure for AI agents

Domain

AI Infrastructure / Distributed Systems

Required skills

distributed systems design, large-scale production system operations, latency optimization, capacity management, system architecture, trade-off analysis, infrastructure ownership

Preferred skills

ML infrastructure, MLOps, LLM serving at scale, self-hosted inference, GPU infrastructure, vLLM, SGLang, post-training infrastructure

Technologies

vLLM, SGLang, GPU infrastructure

Responsibilities

Design inference architecture across self-hosted models and third-party providers; Develop systems for routing, failover, capacity management, and quota; Operate self-hosted inference on GPU infrastructure; Optimize inference performance via speculative decoding and serving-engine tuning; Make architectural decisions for hybrid inference stacks; Support model lifecycle infrastructure.

Seniority

Senior, hands-on IC