Staff HPC Systems Architect
Core
Architect and define scalable compute platforms optimized for AI/ML, simulation, and high-throughput workloads.
Role type
Staff HPC Systems Architect
Builds
Scalable AI/ML compute platforms, rack-level and cluster designs
Domain
High-performance computing, AI infrastructure, hardware architecture
Deliverable
production ML models | infrastructure
Required skills
Large-scale GPU HPC/cloud platform architecture, CPU/GPU/accelerator topology knowledge, high-bandwidth low-latency fabric design (NVLink, InfiniBand, RoCE), system performance tuning, thermal and power optimization, compute lifecycle management, hardware/software boundary navigation, architectural tradeoff analysis
Preferred skills
AI/ML workload performance characteristics, HPC orchestration tools (Slurm, Kubernetes), GPU virtualization, hardware validation, vendor collaboration, compute telemetry, large-scale A/B infrastructure testing
Technologies
NVLink, InfiniBand, RoCE, Slurm, Kubernetes
Responsibilities
Architect scalable compute platforms for AI/ML and simulation; Develop compute system standards and design patterns; Evaluate emerging CPU/GPU/accelerator technologies; Map workload requirements to compute platform capabilities; Define compute platform roadmaps and architectural reference designs; Act as technical lead for new platform introductions and validation; Mentor systems engineers on performance tuning and architectural decisions
Seniority
Staff, hands-on IC with strategic influence
