Staff Software Engineer- Foundation Model Inference
Core
Build LLM infrastructure powering large-scale inference workloads for customers through partner models (OpenAI, Anthropic, Gemini) and self-hosted models (Qwen, GPT-OSS, Llama).
Role type
Staff Software Engineer (Foundation Model Inference).
Builds
Unified platform for serving, scaling, and optimizing frontier models with enterprise-grade reliability.
Domain
Generative AI infrastructure, distributed systems, cloud-native platforms.
Deliverable
production ML models.
Required skills
backend engineering, distributed systems, scalable APIs, cloud-native infrastructure, real-time serving, ML infrastructure, GPU orchestration, service-oriented architecture, deployment pipelines, system observability.
Preferred skills
experience with SageMaker, Vertex AI, Azure ML, contributions to OSS projects like MLflow, Pytorch, Ray, vLLM, SGLang, building developer platforms.
Technologies
OpenAI, Anthropic, Gemini, Qwen, GPT-OSS, Llama, vLLM, SGLang, Ray, PyTorch, MLflow, SageMaker, Vertex AI, Azure ML.
Responsibilities
Build LLM infrastructure for large-scale inference, improve reliability/latency/efficiency of distributed AI workloads, collaborate with platform/infra/ML teams, shape developer experiences for AI on Databricks.
Seniority
Staff, hands-on IC with strategic impact.