Senior Software Engineer II, Applied Training
Core
Build a Kubernetes-native research cluster platform and sandbox infrastructure to enable AI labs to train models efficiently without operational overhead.
Role type
Senior IC distributed systems engineer (ML infrastructure)
Builds
Kubernetes operators, CLI tools, job orchestration schemas, Python SDKs for RL training, and documentation for OSS training frameworks.
Domain
AI infrastructure / Cloud computing / Distributed systems
Deliverable
production ML models | infrastructure
Required skills
Kubernetes (custom controllers, operators, scheduling, CRDs), distributed systems, ML infrastructure, Python, customer-facing system design
Preferred skills
Agentic AI (RL training, agent evaluation), Slurm/Ray experience, container isolation (gVisor, Kata), OSS contributions to Kubernetes/PyTorch
Technologies
Kubernetes, Python, gVisor, Kata, Slurm, Ray, PyTorch
Responsibilities
Design and build complete research cluster experiences including CLI and job configuration schemas; own the Python SDK for sandbox infrastructure enabling RL training runs; write documentation for running OSS training frameworks; collaborate with customers and internal teams to solve researcher productivity problems
Seniority
Senior, hands-on IC