Platform Engineer
Core
Designing and maintaining large-scale GPU compute infrastructure to support reinforcement learning algorithms for a superintelligence discovery mission.
Role type
Senior Platform Engineer (Infrastructure)
Builds
Reliable GPU clusters, developer environments, and resilient cloud infrastructure for AI research.
Domain
Artificial Intelligence / High-Performance Computing / Cloud Infrastructure
Deliverable
infrastructure
Required skills
Kubernetes, GPU scheduling, cloud infrastructure, observability, Python, Rust, system reliability engineering
Preferred skills
Experience with KAI scheduler, Kueue, Tailscale, Workbrew, Google Cloud Platform
Technologies
Kubernetes, KAI scheduler, Kueue, Google Cloud, QuickWit, Grafana, Datadog, Tailscale, Workbrew, Python, Rust
Responsibilities
Manage Kubernetes clusters and deploy internal tools; optimize GPU scheduling at scale; build systems to prevent hardware failures and cluster-scale chaos; maintain cloud infrastructure on Google Cloud and other providers; implement log management and monitoring solutions; develop developer tooling to improve team efficiency; write clean, well-crafted code for infrastructure tools.
Seniority
Senior, hands-on IC