CareerPlanSign in

Member of Technical Staff - ClusterMAX

San Francisco, CA, New York City, NY🌐 Remote💼 Full-time🗓 2026-09-20 → 2026-09-26

Core

Build and run benchmarks to rate GPU cloud providers on performance, security, and cost-efficiency for the AI infrastructure market.

Role type

Senior IC machine-learning infrastructure engineer (benchmarking)

Builds

ClusterMAX™ benchmark suite and evaluation framework for GPU clusters

Domain

Semiconductor industry + AI infrastructure (GPU clouds, hyperscalers, neoclouds)

Deliverable

production ML models | product features

Required skills

GPU cluster operations, Python, shell scripting, distributed systems, security testing, CI/CD pipeline development

Preferred skills

TCO analysis, client-facing technical communication, due diligence

Technologies

Slurm, Kubernetes, InfiniBand, RoCE, NCCL, RCCL, GB300, GB200, B300, B200, H200, MI355X, TPUv7

Responsibilities

Develop next-generation benchmarks for storage IO, bandwidth, collectives, fault-tolerance, and goodput; Deploy and evaluate benchmarks across dozens of GPU clusters; Extend TCO and goodput methodology into automated tests; Author technical research on benchmark results and reliability

Seniority

Senior, hands-on IC

Sourced via dover · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.