MTS, Inference Performance Visibility
Core
Design and develop a sophisticated performance analysis tool for custom ML accelerator hardware to help engineers and customers identify bottlenecks and optimize inference workloads.
Role type
Senior IC systems engineer (performance analysis tooling)
Builds
Performance analysis suite (data collection, processing pipelines, analysis engines, CLI/GUI interfaces)
Domain
Hardware infrastructure for frontier AI inference (ML accelerators, PCIe, system-level tracing)
Deliverable
production ML models | product features
Required skills
C++ or Rust, Python, computer architecture (CPU/GPU/accelerators), memory hierarchies, interconnects (PCIe), low-level performance analysis, profiling, bottleneck identification, performance analysis tools (Nsight, VTune, perf, Tracy, ETW), driver interaction
Preferred skills
Kernel-mode driver development, ML accelerator architectures, compiler internals, PCIe protocol analysis, multi-chip/multi-host systems, firmware/embedded systems, hardware description languages
Technologies
C++, Rust, Python, PCIe, Linux, Windows, NVIDIA Nsight, AMD uProf, Intel VTune, perf, Tracy, ETW
Responsibilities
Lead design and architecture of performance analysis suite; Develop methods to capture performance data from custom ML accelerator hardware; Implement tracing for host-side API calls and system-level events; Design techniques to correlate performance events across CPU, driver, PCIe, and accelerators; Build analysis modules to interpret trace data and identify bottlenecks; Develop visualizations for performance characteristics; Collaborate with hardware architects, firmware, driver, compiler, and ML engineers
Seniority
Senior, hands-on IC
