Software Engineer - Baseten Inference Stack
Core
Build the distributed runtime and orchestration systems that power large-scale LLM inference across the platform, enabling customers to deploy and operate cutting-edge models with high performance and reliability.
Role type
Senior IC platform engineer (LLM inference stack)
Builds
Distributed runtime, orchestration systems, routing, autoscaling, and scheduling for LLM inference
Domain
AI Infrastructure / Distributed Systems / LLM Inference
Deliverable
production ML models
Required skills
distributed systems, backend infrastructure, platform engineering, production system operations, developer experience, system debugging, engineering tradeoffs, end-to-end project ownership
Preferred skills
Kubernetes (operators, custom resources), inference frameworks (Dynamo, vLLM, SGLang, TensorRT-LLM), distributed scheduling, autoscaling, service orchestration, GPU workload operations, observability, CI/CD, open-source contributions
Technologies
Kubernetes, NVIDIA Dynamo, vLLM, SGLang, TensorRT-LLM
Responsibilities
Develop infrastructure and orchestration systems for deploying and managing large-scale distributed LLM inference; Build platform capabilities related to routing, autoscaling, scheduling, observability, and runtime management; Debug complex production systems spanning Kubernetes, distributed runtimes, networking, and GPU workloads; Collaborate with Model Performance engineers to make new inference optimizations broadly available; Define best practices around testing, release automation, and operational excellence; Own projects end-to-end from architecture through deployment and iteration
Seniority
Senior, hands-on IC