Senior Software Engineer, AI Infrastructure
Core
Designing and operating high-performance infrastructure (orchestration, runtime, automation) to support large-scale AI model training on GPU clusters for the global research community.
Role type
Senior Software Engineer, AI Infrastructure
Builds
Beaker job scheduler, execution runtime, and automated cluster management tools
Domain
AI Infrastructure / High-Performance Computing (HPC)
Deliverable
production ML models
Required skills
Go, Python, Linux internals, container runtimes (Docker), distributed systems design, root-cause analysis, system automation
Preferred skills
Kubernetes, Slurm, NCCL, InfiniBand, on-prem storage (WEKA, Ceph), open-source contributions
Technologies
Go, Python, Docker, Kubernetes, Slurm, NCCL, InfiniBand, WEKA, Ceph
Responsibilities
Design and deliver full-stack systems from scheduler to execution runtime; build tooling for cluster health automation; optimize distributed workload performance; define roadmap for compute/networking/storage deployment; mentor team members and drive process improvements.
Seniority
Senior, hands-on IC