CareerPlanGet AI match score →

Staff+ Software Engineer, Inference Runtime

San Francisco, CA💼 Full-time💰 $405,000–$405,000🗓 2026-06-12 → 2026-07-31

Core

Technical lead for the accelerator-agnostic core of Anthropic's inference serving stack, ensuring performance, correctness, and abstractions for GPUs, TPUs, and Trainium.

Role type

Staff+ Software Engineer, Inference Runtime

Builds

The shared, accelerator-agnostic runtime serving Claude to millions of users and enterprise customers

Domain

AI/ML Infrastructure, High-Performance Computing, Accelerator Ecosystems

Deliverable

production ML models

Required skills

Systems engineering, Performance profiling, Latency and throughput optimization, Systems debugging, Accelerator ecosystem depth (CUDA/GPU, TPU, or Trainium), High-performance distributed systems, Engineering metrics definition, Technical alignment across orgs

Preferred skills

ML compiler toolchains (XLA, Triton, NeuronX), Accelerator driver/firmware management, Production validation surface operations, Deterministic/simulation-based testing, CI/CD at scale, Kubernetes-based development, Platform engineering leadership

Technologies

Rust, Python, CUDA, TPU, Trainium, XLA, Triton, NeuronX, Kubernetes

Responsibilities

Set technical direction for the team, owning the architecture and roadmap for the shared runtime, Own and evolve the accelerator-agnostic runtime including hands-on work in Rust and Python, Drive efficient accelerator usage across GPU, TPU, and Trainium, Build the runtime's validation surface around partitioned builds and canary/shadow/rollback mechanisms, Act as a technical counterpart to the central Infrastructure org, Mentor engineers on the team through design review and code review

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗