CareerPlanGet AI match score →

Performance Engineer, Inference Systems

San Francisco, CA💼 Full-time💰 $350,000–$350,000🗓 2026-05-20 → 2026-07-31

Core

Optimizing the throughput, latency, reliability, and correctness of Anthropic's large-scale AI inference fleet serving Claude.

Role type

Senior IC performance engineer (inference systems)

Builds

Observability dashboards, correctness evaluation pipelines, and performance modeling tools for the inference stack.

Domain

AI/ML infrastructure, large-scale distributed systems, GPU/TPU acceleration

Deliverable

production ML models

Required skills

performance engineering, profiling, roofline analysis, latency/throughput optimization, root-cause investigation, Python, data analysis (SQL/pandas), correctness evaluation design

Preferred skills

ML systems experience, GPU/TPU performance concepts, reliability engineering for high-throughput services, model evaluation pipelines, observability for distributed systems

Technologies

Python, SQL, pandas, GPU/TPU, accelerator kernels, model servers, distributed routing, autoscaling

Responsibilities

Run cross-layer performance investigations to size gaps between actual and theoretical performance; own and improve correctness evaluation pipelines; build observability and modeling tools; partner with kernel and serving teams to land optimizations; prioritize opportunities by impact and effort.

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗