Staff+ Software Engineer, Inference Runtime
Core
Technical lead for the accelerator-agnostic core of Anthropic's inference serving stack, ensuring performance, correctness, and abstractions for GPUs, TPUs, and Trainium.
Role type
Staff+ Software Engineer, Inference Runtime
Builds
The shared, accelerator-agnostic runtime serving Claude to millions of users and enterprise customers
Domain
AI/ML Infrastructure, High-Performance Computing, Accelerator Ecosystems
Deliverable
production ML models
Required skills
Systems engineering, Performance profiling, Latency and throughput optimization, Systems debugging, Accelerator ecosystem depth (CUDA/GPU, TPU, or Trainium), High-performance distributed systems, Engineering metrics definition, Technical alignment across orgs
Preferred skills
ML compiler toolchains (XLA, Triton, NeuronX), Accelerator driver/firmware management, Production validation surface operations, Deterministic/simulation-based testing, CI/CD at scale, Kubernetes-based development, Platform engineering leadership
Technologies
Rust, Python, CUDA, TPU, Trainium, XLA, Triton, NeuronX, Kubernetes
Responsibilities
Set technical direction for the team, owning the architecture and roadmap for the shared runtime, Own and evolve the accelerator-agnostic runtime including hands-on work in Rust and Python, Drive efficient accelerator usage across GPU, TPU, and Trainium, Build the runtime's validation surface around partitioned builds and canary/shadow/rollback mechanisms, Act as a technical counterpart to the central Infrastructure org, Mentor engineers on the team through design review and code review