Software Engineer, GPU Infrastructure (HPC)
Core
Building and operating GPU/TPU superclusters and HPC infrastructure to train, evaluate, and serve frontier AI models for developers and enterprises.
Role type
Staff Software Engineer (GPU Infrastructure)
Builds
Kubernetes-based GPU/TPU superclusters across multiple clouds for AI workloads
Domain
High-Performance Computing (HPC) / AI Infrastructure
Deliverable
infrastructure
Required skills
GPU/TPU cluster management, Kubernetes at scale, Python, Go, Linux internals, RDMA networking, distributed training frameworks (JAX, PyTorch, TensorFlow)
Preferred skills
Open-source contributions, observability, infrastructure-as-code (IaC), mentoring
Technologies
Kubernetes, RDMA, NCCL, JAX, PyTorch, TensorFlow
Responsibilities
Deploy and manage Kubernetes-based GPU/TPU superclusters; Optimize infrastructure for cost efficiency and performance using RDMA/NCCL; Troubleshoot complex infrastructure bottlenecks and failures; Design self-service tools for researchers; Collaborate with researchers on emerging ML needs; Advocate for observability and automation; Mentor team members through code reviews and documentation