Staff Software Engineer, Applied Training
Core
Build a Kubernetes-native research cluster platform and sandbox infrastructure to enable AI labs to train models efficiently without operational overhead.
Role type
Staff Software Engineer (ML Infrastructure)
Builds
Kubernetes operators, CLI tools, job orchestration systems, and Python SDKs for distributed training
Domain
AI Infrastructure / Cloud Computing / Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Distributed systems, Kubernetes (operators, CRDs, scheduling), ML training workflows, Python, Customer-facing system design
Preferred skills
Agentic AI/RL training, Slurm/Ray, Container isolation (gVisor/Kata), OSS contributions
Technologies
Kubernetes, Python, Slurm, Ray, gVisor, Kata
Responsibilities
Design and build complete research cluster experiences including CLI and job configuration schemas; Own the Python SDK for sandbox infrastructure enabling isolated agent rollouts; Collaborate with customers to translate research needs into system designs; Write documentation for OSS training frameworks; Contribute to the product roadmap for applied training
Seniority
Staff, hands-on IC with strategic impact