Member of Technical Staff, GPU Infrastructure
Core
Design and own a self-serve compute platform for training jobs and inference services, managing GPU fleets across multiple cloud providers.
Role type
Senior IC GPU infrastructure engineer
Builds
Self-serve compute platform for AI training and inference
Domain
Cloud infrastructure + GPU computing
Deliverable
production ML models | infrastructure
Required skills
Kubernetes (operators, CRDs, multi-cluster), GPU cluster management, distributed systems, Go/Rust/C++, cloud orchestration (AWS, GCP, CoreWeave)
Preferred skills
Inference serving stacks (vLLM, SGLang, TensorRT-LLM), HPC schedulers (Slurm), GPU kernel development (CUDA, Triton), fast interconnects (InfiniBand, RoCE, RDMA)
Technologies
Kubernetes, AWS, CUDA, Grafana, InfiniBand, Prometheus, Rust, vLLM, NodeJS, DevOps
Responsibilities
Design self-serve platform for compute consumption, manage GPU fleet lifecycle and capacity, develop scheduling/placement logic, maintain Kubernetes orchestration layer, strengthen fault tolerance and autoscaling, partner on platform architecture
Seniority
Senior, hands-on IC