Senior Staff System Engineer, GPU Fleet
Core
Senior technical owner for hyperscale GPU compute infrastructure supporting large-scale AI training and inference workloads.
Role type
Sr Staff System Engineer (GPU Fleet)
Builds
Hyperscale GPU compute infrastructure for AI labs, governments, and enterprises
Domain
Cloud Infrastructure / AI Hardware / Datacenter Operations
Deliverable
infrastructure
Required skills
Linux system internals, hardware-intensive infrastructure operations, production-grade automation (Python/Bash), complex multi-layer debugging, fleet architecture design, observability and resilience engineering
Preferred skills
Large-scale GPU fleet operations, modern GPU ecosystems (CUDA, NCCL), high-speed interconnects (NVLink, InfiniBand, RDMA), fleet-wide initiative ownership, hardware vendor partnership, failure mode analysis at scale
Technologies
Linux, Python, Bash, CUDA, NCCL, NVLink, InfiniBand, RDMA
Responsibilities
Define end-to-end technical architecture for GPU fleets including hardware selection and networking; Drive fleet-level reliability, availability, and performance objectives; Design and build large-scale automation for provisioning, health validation, and remediation; Act as senior escalation point for critical production incidents; Provide technical mentorship and influence long-term infrastructure roadmap
Seniority
Sr Staff, hands-on IC with principal-level scope