GPU Cluster Architect
Core
Design and architect next-generation AI infrastructure for tens of thousands of GPUs across multiple data center sites, ensuring scalability, performance, and reliability for modern AI workloads.
Role type
Senior IC GPU Cluster Architect
Builds
Full-stack AI cloud platform supporting data and model training through production deployment
Domain
Cloud infrastructure, High-Performance Computing (HPC), AI/ML systems
Deliverable
infrastructure
Required skills
GPU cluster topology design, HPC interconnects (InfiniBand, RoCE), performance modeling for AI/ML workloads, systems architecture, networking, hardware reliability, automation scripting (Python, Go)
Preferred skills
Experience with NVIDIA/AMD GPU architectures, low-latency network design, storage integration for training datasets
Technologies
InfiniBand, Ethernet, RoCEv2, Python, Go
Responsibilities
Architect scalable GPU cluster topologies including compute nodes, interconnects, storage, and control planes; Analyze AI/ML workloads to inform design tradeoffs across latency, bandwidth, and GPU density; Align with network architect to validate low-latency, high-throughput interconnects at POD and DC scale; Work with storage teams to optimize performance for training datasets and checkpointing; Analyze monitoring signals to detect design flows; Partner with site reliability, networking, storage, and DC engineering teams to operationalize and scale architecture
Seniority
Senior, hands-on IC