Member of Technical Staff - ClusterMAX
Core
Build and run benchmarks to rate GPU cloud providers on performance, security, and cost-efficiency for the AI infrastructure market.
Role type
Senior IC machine-learning infrastructure engineer (benchmarking)
Builds
ClusterMAX™ benchmark suite and evaluation framework for GPU clusters
Domain
Semiconductor industry + AI infrastructure (GPU clouds, hyperscalers, neoclouds)
Deliverable
production ML models | product features
Required skills
GPU cluster operations, Python, shell scripting, distributed systems, security testing, CI/CD pipeline development
Preferred skills
TCO analysis, client-facing technical communication, due diligence
Technologies
Slurm, Kubernetes, InfiniBand, RoCE, NCCL, RCCL, GB300, GB200, B300, B200, H200, MI355X, TPUv7
Responsibilities
Develop next-generation benchmarks for storage IO, bandwidth, collectives, fault-tolerance, and goodput; Deploy and evaluate benchmarks across dozens of GPU clusters; Extend TCO and goodput methodology into automated tests; Author technical research on benchmark results and reliability
Seniority
Senior, hands-on IC