Performance Architect, AI HW
Core
Modeling, analyzing, and optimizing real AI workloads on the Tensix architecture to shape future hardware features and ensure design decisions deliver measurable performance gains.
Role type
AI Performance Architect (Hardware/Software Co-design)
Builds
Next-generation AI compute fabrics and high-performance RISC-V CPU systems
Domain
AI Hardware Acceleration / Semiconductor Architecture
Deliverable
production ML models | infrastructure
Required skills
C++, Python, system-level performance analysis, PPA (Performance/Power/Area) studies, heterogeneous compute systems, RTL collaboration, deep learning workload behavior
Preferred skills
multi-chip/distributed performance analysis, compiler and runtime layer knowledge, custom AI accelerator design
Technologies
C++, Python, RISC-V, Tensix architecture, RTL, deep learning frameworks
Responsibilities
Benchmark and analyze complex AI workloads across single and multi-node hardware configurations; Develop and maintain performance models, simulators, and micro-benchmark suites; Conduct detailed PPA studies to assess design tradeoffs; Collaborate with RTL, Compiler, and Runtime teams to instrument and correlate performance models with silicon results
Seniority
Individual Contributor (various levels assessed during interview)