CareerPlanSign in

Staff Software Engineer- Foundation Model Inference

Mountain View, California💼 Full-time💰 $190,000–$190,000🗓 2026-07-24 → 2026-09-25

Core

Build LLM infrastructure powering large-scale inference workloads for customers through partner models (OpenAI, Anthropic, Gemini) and self-hosted models (Qwen, GPT-OSS, Llama).

Role type

Staff Software Engineer (Foundation Model Inference).

Builds

Unified platform for serving, scaling, and optimizing frontier models with enterprise-grade reliability.

Domain

Generative AI infrastructure, distributed systems, cloud-native platforms.

Deliverable

production ML models.

Required skills

backend engineering, distributed systems, scalable APIs, cloud-native infrastructure, real-time serving, ML infrastructure, GPU orchestration, service-oriented architecture, deployment pipelines, system observability.

Preferred skills

experience with SageMaker, Vertex AI, Azure ML, contributions to OSS projects like MLflow, Pytorch, Ray, vLLM, SGLang, building developer platforms.

Technologies

OpenAI, Anthropic, Gemini, Qwen, GPT-OSS, Llama, vLLM, SGLang, Ray, PyTorch, MLflow, SageMaker, Vertex AI, Azure ML.

Responsibilities

Build LLM infrastructure for large-scale inference, improve reliability/latency/efficiency of distributed AI workloads, collaborate with platform/infra/ML teams, shape developer experiences for AI on Databricks.

Seniority

Staff, hands-on IC with strategic impact.

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.