Member of Technical Staff - Compute Platform
Core
Build and maintain tools for automatic remediation, topology-aware scheduling, capacity planning, and rapid hardware debugging for large, multi-GPU fleets on a K8s-based multi-cloud platform.
Role type
Senior IC systems engineer (compute platform)
Builds
Cluster management stack, monitoring & observability tools, and infrastructure for next-generation GPU deployments
Domain
Cloud infrastructure + GPU systems
Deliverable
infrastructure
Required skills
Systems-level engineering, K8s architecture, GPU hardware knowledge, Cloud storage management, Performance benchmarking, Topology-aware scheduling
Preferred skills
NCCL familiarity, Petabyte-scale data replication, GPU-to-GPU network optimization
Technologies
Kubernetes, NCCL, VAST
Responsibilities
Design and iterate on cluster management stack, Implement comprehensive cluster-wide monitoring, Prepare infrastructure for next-generation GPU deployments, Manage multi-cloud storage and data replication
Seniority
Senior, hands-on IC