Software Engineer, Inference
Core
Design and build large-scale inference systems for AI agents, managing routing, capacity, and optimization across self-hosted models and third-party providers.
Role type
Senior IC distributed systems engineer (AI inference)
Builds
High-performance inference serving stack, routing layers, and GPU infrastructure for AI agents
Domain
AI Infrastructure / Distributed Systems
Deliverable
production ML models (via careerplan.io/jobs/b8666a78-fcb1-47bb-ad07-5edb6cd3ad71-software-engineer-inference-at-sierra)
Required skills
distributed systems design, large-scale production system operations, latency optimization, capacity management, system architecture, trade-off analysis, infrastructure ownership
Preferred skills
ML infrastructure, MLOps, LLM serving at scale, self-hosted inference, GPU infrastructure, vLLM, SGLang, post-training infrastructure
Technologies
vLLM, SGLang, GPU infrastructure
Responsibilities
Design inference architecture across self-hosted models and third-party providers; Develop systems for routing, failover, capacity management, and quota; Operate self-hosted inference on GPU infrastructure; Optimize inference performance via speculative decoding and serving-engine tuning; Make architectural decisions for hybrid inference stacks; Support model lifecycle infrastructure.
Seniority
Senior, hands-on IC
