Member of Technical Staff - GPU Infrastructure
Core
Design and deploy large-scale GPU cluster architectures for customer AI workloads, serving as the primary technical escalation point for infrastructure issues.
Role type
Senior IC GPU Infrastructure Engineer (Customer-Facing)
Builds
High-performance GPU clusters (100-10,000+ GPUs) for LLM training, inference, and HPC
Domain
AI Infrastructure / High-Performance Computing
Deliverable
production ML models | infrastructure
Required skills
GPU cluster architecture, SLURM, Kubernetes, InfiniBand configuration, NVIDIA CUDA ecosystem, Linux kernel tuning, Python, Bash, systems programming
Preferred skills
1000+ GPU deployments, NVIDIA DGX/HGX/SuperPOD certification, distributed training frameworks (PyTorch FSDP, DeepSpeed), AMD MI300/Intel Gaudi experience
Technologies
SLURM, Kubernetes, InfiniBand, RoCE, NVLink, Lustre, BeeGFS, GPFS, Ansible, Terraform, Docker, Containerd, Enroot
Responsibilities
Partner with clients to design optimal GPU cluster architectures; Deploy and configure orchestration systems and high-performance networking; Serve as primary technical escalation point for customer infrastructure issues; Diagnose and resolve complex hardware, driver, and networking problems; Implement monitoring, alerting, and automated remediation systems.
Seniority
Senior, hands-on IC