CareerPlanSign in

Head of Inference Performance Visibility

San Jose💼 Full-time🗓 2026-02-19 → 2026-09-25

Core

Define performance metrics and abstractions connecting raw hardware signals to distributed AI workload context for next-generation inference systems.

Role type

Head of Inference Performance Visibility (Senior IC/Technical Lead)

Builds

Scalable performance reasoning systems, cross-layer event correlation frameworks, and cluster-scale analysis tools for distributed inference.

Domain

AI Infrastructure / High-Performance Computing / Custom Silicon

Deliverable

production ML models | infrastructure

Required skills

Systems programming (C++/Rust), distributed tracing, hardware counter design, time synchronization, cross-layer event correlation, performance modeling, observability platform design, low-level infrastructure development.

Preferred skills

Experience with custom silicon, supercomputing software, compiler stacks, runtime libraries, large AI training clusters, multi-device systems.

Technologies

C++, Rust, distributed systems, hardware counters, telemetry frameworks.

Responsibilities

Define architectural approaches for collecting and structuring telemetry across CPUs, drivers, and accelerators; design scalable models for correlating performance events across device and host boundaries; implement time synchronization and trace-alignment strategies; build tools to identify bottlenecks in multi-accelerator workloads; contribute to analysis engines transforming raw telemetry into insights.

Seniority

Senior, hands-on IC with leadership responsibilities

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.