Staff Software Engineer, Foundation Model API
Core
Build LLM infrastructure powering large-scale inference workloads for customers through partner models (OpenAI, Anthropic, Gemini) and self-hosted models (Qwen, GPT-OSS, Llama).
Role type
Staff Software Engineer (Foundation Model API Infrastructure).
Builds
Unified serving layer for large language models across real-time and batch inference.
Domain
AI Infrastructure / Large Language Model Serving.
Deliverable
production ML models.
Required skills
Backend engineering, distributed systems, scalable APIs, cloud-native infrastructure, real-time serving, ML infrastructure, GPU orchestration, service-oriented architecture, deployment pipelines, system observability, Scala, Go, Python.
Preferred skills
Experience with SageMaker, Vertex AI, Azure ML, building products supporting AI workflows.
Responsibilities
Build LLM infrastructure for large-scale inference, shape product direction via customer engagement, improve reliability/latency/efficiency of distributed AI workloads, collaborate with platform/infra/ML teams, shape developer experiences for AI on Databricks.
Seniority
Staff, high-agency IC with product strategy influence.