Inference Systems Performance Architect
Core
Define and drive technical strategy for end-to-end performance of large-scale LLM inference, including workload capture, benchmarking, modeling, and simulation to optimize serving and inform system planning.
Role type
Senior IC Inference Systems Performance Architect
Builds
Heterogeneous, disaggregated inference systems (GPU on prefill, RDU on decode) and performance tooling for capacity planning
Domain
Generative AI / Large Language Models / High-Performance Computing
Deliverable
production ML models
Required skills
End-to-end performance analysis of distributed systems, realistic workload generation, performance modeling and simulation, cross-functional technical leadership, customer-facing technical communication, scoping high-ambiguity work
Preferred skills
LLM inference serving (continuous batching, prompt/KV caching, prefill/decode disaggregation), inference simulation frameworks, public technical voice
Technologies
SN40L chip, SambaNova Suite, open-source LLMs, distributed inference pipelines
Responsibilities
Define technical strategy for inference-systems performance; build workload-capture and agentic-benchmarking capabilities; own performance-modeling and simulation practice; drive end-to-end profiling tooling; serve as senior technical voice across model-optimization, systems, hardware, and product; mentor principal and senior engineers; resolve novel challenges spanning organizational boundaries
Seniority
Senior, hands-on IC with strategic scope