Senior Software Engineer, GPU Infrastructure
Core
Build, deploy, and operate Kubernetes-based GPU/TPU superclusters to optimize AI/ML training performance, reliability, and cost.
Role type
Senior IC GPU infrastructure engineer
Builds
Scalable GPU/TPU clusters and self-service interfaces for distributed training workloads
Domain
Cloud infrastructure + AI/ML high-performance computing
Deliverable
infrastructure
Required skills
Kubernetes cluster management at scale, GPU/TPU cluster operations, distributed training frameworks (JAX, PyTorch), Python, Go, Linux internals, RDMA networking, performance optimization
Preferred skills
Open-source contribution experience
Technologies
Kubernetes, RDMA, NCCL, JAX, PyTorch
Responsibilities
Build and operate GPU/TPU superclusters across multiple clouds; optimize infrastructure for AI/ML training; diagnose and resolve infrastructure bottlenecks and failures; create self-service interfaces for researchers; collaborate with AI researchers and ML engineers; participate in 24x7 on-call rotation
Seniority
Senior, hands-on IC