CareerPlanSign in

Inference Systems Performance Architect

San Jose, CA💼 Full-time💰 $245,000–$245,000🗓 2026-08-16 → 2026-09-25

Core

Define and drive technical strategy for end-to-end performance of large-scale LLM inference, including workload capture, benchmarking, modeling, and simulation to optimize serving and inform system planning.

Role type

Senior IC Inference Systems Performance Architect

Builds

Heterogeneous, disaggregated inference systems (GPU on prefill, RDU on decode) and performance tooling for capacity planning

Domain

Generative AI / Large Language Models / High-Performance Computing

Deliverable

production ML models

Required skills

End-to-end performance analysis of distributed systems, realistic workload generation, performance modeling and simulation, cross-functional technical leadership, customer-facing technical communication, scoping high-ambiguity work

Preferred skills

LLM inference serving (continuous batching, prompt/KV caching, prefill/decode disaggregation), inference simulation frameworks, public technical voice

Technologies

SN40L chip, SambaNova Suite, open-source LLMs, distributed inference pipelines

Responsibilities

Define technical strategy for inference-systems performance; build workload-capture and agentic-benchmarking capabilities; own performance-modeling and simulation practice; drive end-to-end profiling tooling; serve as senior technical voice across model-optimization, systems, hardware, and product; mentor principal and senior engineers; resolve novel challenges spanning organizational boundaries

Seniority

Senior, hands-on IC with strategic scope

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.