CareerPlanGet AI match score →

Staff Software Engineer, Inference Infrastructure

San Francisco🌐 Remote💼 Full-time🗓 2026-01-12 → 2026-07-31

Core

Designing, deploying, and operating high-performance, scalable machine learning systems for serving large language models via API endpoints.

Role type

Staff Software Engineer, Inference Infrastructure

Builds

AI platform delivering large language models through easy-to-use API endpoints

Domain

Artificial Intelligence / Large Language Models / Cloud Infrastructure

Deliverable

production ML models

Required skills

distributed systems design, Kubernetes orchestration, GPU workload management, multi-cloud infrastructure (GCP/Azure/AWS/OCI), Linux system administration, compute/storage/network resource optimization, high-performance server development (Golang/C++)

Preferred skills

familiarity with accelerator characteristics (GPUs/TPUs/custom accelerators), customer-facing deployment customization

Technologies

Kubernetes, GCP, Azure, AWS, OCI, Golang, C++, Linux

Responsibilities

Develop and deploy optimized NLP models to production environments, operate AI platform for low latency/high throughput, interface with customers for customized deployments, troubleshoot complex Linux-based computing environments, manage compute/storage/network resources and costs

Seniority

Staff, hands-on IC with strategic impact

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗