Senior Software Engineer II, Applied Training
Core
Build a Kubernetes-native research cluster platform and sandbox infrastructure to enable AI labs to train models efficiently on CoreWeave's GPU cloud.
Role type
Senior IC distributed systems engineer (ML infrastructure)
Builds
Production-grade research cluster platform, CLI tools, Kubernetes operators, Python SDKs for RL training, and sandbox environments for agent rollouts.
Domain
Cloud infrastructure for AI/ML training and distributed systems
Deliverable
production ML models | product features
Required skills
Distributed systems, Kubernetes (operators, CRDs, scheduling), ML infrastructure, Python, customer-facing system design, documentation
Preferred skills
Agentic AI/RL training, Slurm/Ray, container isolation (gVisor/Kata), OSS contributions to Kubernetes/PyTorch
Responsibilities
Design and build complete research cluster experiences including CLI and job orchestration; own the Python SDK for sandbox infrastructure; collaborate with customers and internal teams to solve scale challenges; write documentation for OSS training frameworks.
Seniority
Senior, hands-on IC