CareerPlanSign in

Performance Engineer, Inference Engine

San Francisco, CA💼 Full-time💰 $350,000–$350,000🗓 2026-09-09 → 2026-09-25

Core

Building and optimizing Anthropic's in-house inference engine to manage the token path between accelerator kernels and the routing layer, serving Claude to millions of users.

Role type

Senior IC performance engineer (LLM inference)

Builds

Inference engine software for accelerator platforms

Domain

AI/ML infrastructure, High-performance systems, Distributed systems

Deliverable

production ML models

Required skills

Systems programming (Rust, C++), LLM inference mental model, Performance analysis and profiling, Hardware awareness (FLOPs, HBM, PCIe, RDMA), Model state management, Batching and memory management

Preferred skills

GPU/Accelerator programming, OS internals, Transformer architecture, Allocator/cache/scheduler design, High-bandwidth transport, Determinism and reproducibility

Technologies

Rust, C++, Accelerators, Cloud platforms

Responsibilities

Improve throughput, cost, reliability, and latency across accelerator and cloud platforms; Keep device utilization high; Reuse model state instead of recomputing; Build observability to measure and model performance gaps; Ensure model quality and safety on every token

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.