CareerPlanSign in

Staff GPU Inference SDET

Sunnyvale, CA💼 Full-time🗓 2026-09-09 → 2026-09-25

Core

Founding quality, reliability, and validation lead for a new GPU Inference Development team, designing and scaling end-to-end release qualification and automated test ecosystems for GPU inference stacks and rack-scale compute fleets.

Role type

Staff SDET (GPU Inference Infrastructure)

Builds

Automated test suites, release qualification pipelines, benchmarking tools, and fault-injection suites for multi-node GPU clusters and LLM serving engines.

Domain

AI Infrastructure / GPU Computing / Distributed Systems

Deliverable

production ML models | infrastructure

Required skills

SDET/Infrastructure Quality Lead experience, multi-node GPU cluster provisioning and validation, LLM serving engine architecture, Python automation framework design, container orchestration (Kubernetes/Slurm/Ray), high-performance networking (InfiniBand/RoCE/NCCL), root-cause analysis in distributed systems, chaos engineering, CI/CD integration.

Preferred skills

AMD (ROCm/HIP) or NVIDIA software stack experience, workload replay tooling, MLPerf Inference benchmarking, low-level kernel profiling (PyTorch Profiler/NVTX), C++.

Technologies

Kubernetes, Slurm, Ray, InfiniBand, RoCE, NCCL, Prometheus, Grafana, Python, C++.

Responsibilities

Design and implement automated test automation frameworks and regression gates for the GPU inference stack; benchmark and stress-test distributed LLM serving frameworks; build automated workload replay and benchmarking tools; engineer chaos engineering and fault-injection suites; integrate automated test pipelines with telemetry tools.

Seniority

Staff, hands-on IC with founding team responsibilities

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.