Staff GPU Inference SDET
Core
Founding quality, reliability, and validation lead for a new GPU Inference Development team, designing and scaling end-to-end release qualification and automated test ecosystems for GPU inference stacks and rack-scale compute fleets.
Role type
Staff SDET (GPU Inference Infrastructure)
Builds
Automated test suites, release qualification pipelines, benchmarking tools, and fault-injection suites for multi-node GPU clusters and LLM serving engines.
Domain
AI Infrastructure / GPU Computing / Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
SDET/Infrastructure Quality Lead experience, multi-node GPU cluster provisioning and validation, LLM serving engine architecture, Python automation framework design, container orchestration (Kubernetes/Slurm/Ray), high-performance networking (InfiniBand/RoCE/NCCL), root-cause analysis in distributed systems, chaos engineering, CI/CD integration.
Preferred skills
AMD (ROCm/HIP) or NVIDIA software stack experience, workload replay tooling, MLPerf Inference benchmarking, low-level kernel profiling (PyTorch Profiler/NVTX), C++.
Technologies
Kubernetes, Slurm, Ray, InfiniBand, RoCE, NCCL, Prometheus, Grafana, Python, C++.
Responsibilities
Design and implement automated test automation frameworks and regression gates for the GPU inference stack; benchmark and stress-test distributed LLM serving frameworks; build automated workload replay and benchmarking tools; engineer chaos engineering and fault-injection suites; integrate automated test pipelines with telemetry tools.
Seniority
Staff, hands-on IC with founding team responsibilities