Member of Technical Staff (Software Engineer, Inference & Training Platform)
Core
Build a unified, self-serve platform to manage GPU clusters for training and inference workloads, hiding infrastructure complexity from engineers.
Role type
Senior IC platform engineer (GPU infrastructure & orchestration)
Builds
Self-serve compute platform for training jobs and production inference services
Domain
Cloud infrastructure, GPU clusters, distributed systems
Deliverable
infrastructure
Required skills
Kubernetes (operators, CRDs, multi-cluster federation), GPU cluster management (NVIDIA, CUDA, InfiniBand/RoCE), multi-cloud orchestration, distributed systems (scheduling, resource allocation, fault tolerance), systems programming (Go, Rust, C++)
Preferred skills
Inference serving stacks (vLLM, SGLang, TensorRT-LLM), HPC schedulers (Slurm), GPU kernel development (CUDA, Triton), high-speed interconnects (RDMA), ML observability (Prometheus, Grafana, Weights & Biases)
Technologies
Kubernetes, NVIDIA GPUs, CUDA, InfiniBand, RoCE, CoreWeave, AWS, GCP, Go, Rust, C++, vLLM, SGLang, TensorRT-LLM, Slurm, Prometheus, Grafana, Weights & Biases
Responsibilities
Design and own systems for launching training jobs and operating inference services; manage GPU fleet provisioning, lifecycle, reliability, and capacity across providers; build scheduling logic to optimize capacity usage; maintain health of distributed training jobs and production inference services; write operators and CRDs for GPU orchestration across multiple clusters; implement fault tolerance, autoscaling, and observability; define technical direction and architecture for the platform
Seniority
Senior, hands-on IC with architectural ownership