CareerPlanSign in

Software Engineer - Baseten Inference Stack

San Francisco💼 Full-time🗓 2026-06-02 → 2026-09-25

Core

Build the distributed runtime and orchestration systems that power large-scale LLM inference across the platform, enabling customers to deploy and operate cutting-edge models with high performance and reliability.

Role type

Senior IC platform engineer (LLM inference stack)

Builds

Distributed runtime, orchestration systems, routing, autoscaling, and scheduling for LLM inference

Domain

AI Infrastructure / Distributed Systems / LLM Inference

Deliverable

production ML models

Required skills

distributed systems, backend infrastructure, platform engineering, production system operations, developer experience, system debugging, engineering tradeoffs, end-to-end project ownership

Preferred skills

Kubernetes (operators, custom resources), inference frameworks (Dynamo, vLLM, SGLang, TensorRT-LLM), distributed scheduling, autoscaling, service orchestration, GPU workload operations, observability, CI/CD, open-source contributions

Technologies

Kubernetes, NVIDIA Dynamo, vLLM, SGLang, TensorRT-LLM

Responsibilities

Develop infrastructure and orchestration systems for deploying and managing large-scale distributed LLM inference; Build platform capabilities related to routing, autoscaling, scheduling, observability, and runtime management; Debug complex production systems spanning Kubernetes, distributed runtimes, networking, and GPU workloads; Collaborate with Model Performance engineers to make new inference optimizations broadly available; Define best practices around testing, release automation, and operational excellence; Own projects end-to-end from architecture through deployment and iteration

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.