CareerPlanGet AI match score →

Principal Engineer - Perf and Benchmarking

Bellevue, WA💼 Full-time💰 $206,000–$206,000🗓 2026-07-01 → 2026-07-31

Core

Technical lead for planet-scale performance data ingestion, storage, and analysis, plus driving industry-leading MLPerf benchmarking publications.

Role type

Principal Engineer (Performance & Benchmarking)

Builds

Planet-scale performance data warehouse; MLPerf (Training & Inference) submissions; Kubernetes-native benchmarking service.

Domain

AI Infrastructure / High-Performance Computing / Cloud Data Warehousing

Deliverable

production ML models | dashboards & analysis | infrastructure

Required skills

Distributed systems architecture, HPC/cloud services, GPU performance tuning (CUDA, NCCL, RDMA), Model-server stacks (Triton, vLLM, TensorRT-LLM), Distributed training frameworks (PyTorch FSDP, DeepSpeed, Megatron-LM), Kubernetes, ML control planes, Time-series databases, Log-structured merge trees (LSM), Supply-chain integrity (SBOMs, Cosign)

Preferred skills

MLPerf submission experience, OSS contributions, Multi-region fleet benchmarking, Publications on ML performance

Technologies

Kubernetes, SUNK, Kueue, Kubeflow, Prometheus, Grafana, OpenTelemetry, NVIDIA (Megatron-LM, TensorRT-LLM, DGX cloud), PyTorch, DeepSpeed, vLLM, Triton, KServe, ONNX Runtime

Responsibilities

Define multi-year benchmarking strategy and roadmap; Lead end-to-end MLPerf submissions; Design and maintain Kubernetes-native benchmarking service; Build CI/CD pipelines and observability integrations; Partner with NVIDIA and OSS projects for optimizations.

Seniority

Principal, strategy & mentorship

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗