Member of Technical Staff (Infrastructure Engineer, Compute Infrastructure)
Core
Build and operate the operational layer for GPU clusters, scheduling, networking, and observability to support frontier AI research and agent creation.
Role type
Senior Infrastructure Engineer (Compute)
Builds
GPU clusters, Kubernetes/Linux/cloud stacks, and agent-driven automation for cluster lifecycle.
Domain
AI Research Infrastructure / High-Performance Computing
Deliverable
production ML models | infrastructure
Required skills
Kubernetes internals, cluster provisioning, Linux systems, networking, container orchestration, systems thinking, operational excellence, agent-driven automation
Preferred skills
GPU/CUDA expertise, cloud cluster networking (VPC, BGP, CNI, eBPF), Infrastructure-as-Code (Terraform), leading multi-quarter initiatives
Technologies
Kubernetes, Linux, Cloud environments, Terraform, GPU clusters
Responsibilities
Run and evolve GPU clusters for scheduling and performance; Scale Kubernetes and networking stacks; Establish incident response and on-call health; Build automation for cluster lifecycle; Partner with research teams to optimize the platform.
Seniority
Senior, hands-on IC