CareerPlanSign in

MTS, Inference Performance Visibility

San Jose💼 Full-time🗓 2026-01-18 → 2026-09-25

Core

Design and develop a sophisticated performance analysis tool for custom ML accelerator hardware to help engineers and customers identify bottlenecks and optimize inference workloads.

Role type

Senior IC systems engineer (performance analysis tooling)

Builds

Performance analysis suite (data collection, processing pipelines, analysis engines, CLI/GUI interfaces)

Domain

Hardware infrastructure for frontier AI inference (ML accelerators, PCIe, system-level tracing)

Deliverable

production ML models | product features

Required skills

C++ or Rust, Python, computer architecture (CPU/GPU/accelerators), memory hierarchies, interconnects (PCIe), low-level performance analysis, profiling, bottleneck identification, performance analysis tools (Nsight, VTune, perf, Tracy, ETW), driver interaction

Preferred skills

Kernel-mode driver development, ML accelerator architectures, compiler internals, PCIe protocol analysis, multi-chip/multi-host systems, firmware/embedded systems, hardware description languages

Technologies

C++, Rust, Python, PCIe, Linux, Windows, NVIDIA Nsight, AMD uProf, Intel VTune, perf, Tracy, ETW

Responsibilities

Lead design and architecture of performance analysis suite; Develop methods to capture performance data from custom ML accelerator hardware; Implement tracing for host-side API calls and system-level events; Design techniques to correlate performance events across CPU, driver, PCIe, and accelerators; Build analysis modules to interpret trace data and identify bottlenecks; Develop visualizations for performance characteristics; Collaborate with hardware architects, firmware, driver, compiler, and ML engineers

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.