Software Engineer- Inference Platform
Core
Build the distributed runtime and orchestration systems that power large-scale LLM inference, ensuring models are fast, reliable, and cost-efficient for customers.
Role type
Senior IC distributed systems engineer (inference platform)
Builds
Distributed runtime for LLM inference, Model APIs, and multi-cloud capacity management
Domain
AI Infrastructure / Distributed Systems / LLM Serving (via careerplan.io/jobs/14c4a663-0b1f-4c11-93ff-1359741ee456-software-engineer-inference-platform-at-baseten)
Deliverable
production ML models
Required skills
distributed systems, backend infrastructure, large-scale APIs, low-latency services, infrastructure profiling, SLO management, Kubernetes, observability, release automation
Preferred skills
LLM inference engines (vLLM, SGLang, TensorRT-LLM, TGI, Dynamo), GPU workloads, service meshes, API gateways, distributed scheduling, open-source contributions
Technologies
Kubernetes, vLLM, SGLang, TensorRT-LLM, TGI, Dynamo
Responsibilities
Build infrastructure and orchestration systems for deploying and running distributed LLM inference; Design and operate Model APIs with advanced capabilities like structured outputs and tool calling; Implement platform fundamentals including API versioning, metering, and authentication; Instrument deep observability and build benchmarks for speed and reliability; Debug and harden production systems spanning networking and GPU workloads; Partner with performance teams to optimize inference; Own projects end-to-end from architecture to deployment.
Seniority
Senior, hands-on IC
