Member of Technical Staff - GPU Infrastructure Engineer
Core
Hands-on software engineer ensuring reliability and efficiency of GPU clusters for foundation model training and research.
Role type
Senior IC GPU Infrastructure Engineer
Builds
Production-quality infrastructure tooling, automation, and monitoring for GPU compute environments
Domain
AI Infrastructure / Distributed Systems / HPC
Deliverable
infrastructure
Required skills
Linux, networking, storage, distributed systems, production software engineering, cluster operations, automation
Preferred skills
SLURM, Kubernetes, Ray, Hadoop, GPU/HPC training infrastructure, cloud providers, infrastructure control planes
Technologies
Linux, SLURM, Kubernetes, Ray, Hadoop
Responsibilities
Own reliability and operation of GPU clusters, debug issues across compute/storage/networking/schedulers, improve resource utilization via tooling, onboard/migrate workloads across providers, build monitoring and platform abstractions, contribute to long-term infrastructure architecture
Seniority
Senior, hands-on IC