CareerPlanSign in

Software Engineer - Model Products

San Francisco💼 Full-time🗓 2025-10-11 → 2026-09-25

Core

Design, build, and operate Model APIs and inference runtimes to ensure AI models are fast, reliable, and cost-efficient for developers.

Role type

Senior IC backend engineer (LLM inference systems)

Builds

High-performance Model APIs, inference runtimes, and benchmarking frameworks for open-source LLMs

Domain

AI infrastructure / LLM serving / Distributed systems

Deliverable

production ML models

Required skills

distributed systems, large-scale APIs, low-latency backend services, performance profiling, tracing, capacity planning, SLO management, debugging complex systems, API versioning, usage metering, quotas, authentication

Preferred skills

LLM runtimes (vLLM, SGLang, TensorRT-LLM), Kubernetes, service meshes, API gateways, distributed scheduling, open-source APIs

Technologies

TensorRT-LLM, CUDA, JSON mode, grammar-constrained generation, tool/function calling, multi-modal serving, speculative decoding, guided generation, KV-cache reuse

Responsibilities

Design and operate Model APIs with advanced inference capabilities; Profile and optimize TensorRT-LLM kernels and CUDA operators; Productionize performance improvements across runtimes; Build comprehensive benchmarking frameworks; Instrument deep observability and build repeatable benchmarks; Implement platform fundamentals like API versioning and metering; Collaborate with teams to deliver robust model serving experiences

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.