Cluster Architect
Core
Lead the low-level design and deployment of high-density AI supercomputer systems, translating high-level customer requirements into operational hardware and infrastructure blueprints.
Role type
Senior IC Cluster Architect (HPC/AI Infrastructure)
Builds
High-performance computing (HPC) and artificial intelligence (AI) supercomputer systems
Domain
Datacenter engineering, GPU systems, high-performance networking, and AI infrastructure
Deliverable
production ML models | infrastructure
Required skills
Datacenter engineering (mechanical, electrical, plumbing), GPU systems architecture, high-performance networking (InfiniBand), system-level OS and Linux kernel drivers, infrastructure design documentation, vendor technical evaluation
Preferred skills
HPC/Telco/hyperscale supercomputing experience, TCP/IP stack and networking fundamentals, cloud orchestration (Kubernetes, Slurm), NVIDIA SDKs and networking technologies, C/C++ programming
Technologies
InfiniBand, Linux, Kubernetes, Docker, Slurm, NVIDIA CUDA, RoCE, DPU, ARM CPU
Responsibilities
Lead low-level design of compute, interconnect, storage, and management for AI supercomputers; Evaluate tradeoffs between performance, power, cooling, and cost during procurement; Drive creation of rack layouts, data center floor plans, power/cooling architecture, and Bill of Materials; Participate in Architecture Review Board to ensure design alignment; Produce architecture diagrams, configuration documentation, and operational runbooks for handover to operations.
Seniority
Senior, hands-on IC
