Performance Engineer, Inference Engine
Core
Building and optimizing Anthropic's in-house inference engine to manage the token path between accelerator kernels and the routing layer, serving Claude to millions of users.
Role type
Senior IC performance engineer (LLM inference)
Builds
Inference engine software for accelerator platforms
Domain
AI/ML infrastructure, High-performance systems, Distributed systems
Deliverable
production ML models
Required skills
Systems programming (Rust, C++), LLM inference mental model, Performance analysis and profiling, Hardware awareness (FLOPs, HBM, PCIe, RDMA), Model state management, Batching and memory management
Preferred skills
GPU/Accelerator programming, OS internals, Transformer architecture, Allocator/cache/scheduler design, High-bandwidth transport, Determinism and reproducibility
Technologies
Rust, C++, Accelerators, Cloud platforms
Responsibilities
Improve throughput, cost, reliability, and latency across accelerator and cloud platforms; Keep device utilization high; Reuse model state instead of recomputing; Build observability to measure and model performance gaps; Ensure model quality and safety on every token
Seniority
Senior, hands-on IC