Software Engineer, ML Infrastructure
Core
Building large-scale compute, storage, and software infrastructure to support training the world's best agentic coding model.
Role type
Senior IC ML Infrastructure Engineer
Builds
High-performance GPU infrastructure, training frameworks, and workload scheduling systems for agentic coding models
Domain
AI/ML Infrastructure, Distributed Systems, Cloud & Bare Metal
Deliverable
infrastructure
Required skills
Python, TypeScript, Rust, Golang, distributed storage, Linux systems, Kubernetes, infrastructure-as-code, workload scheduling, data movement systems
Preferred skills
Nvidia GPU operations (Infiniband/RoCE), Ray, Slurm
Technologies
Kubernetes, Linux, Python, TypeScript, Rust, Golang, Ray, Slurm
Responsibilities
Collaborate with ML researchers to improve training throughput and reliability; plan and build cutting-edge GPU infrastructure with OEMs and cloud providers; improve compute density and scalability for large RL workloads; create software to automate GPU cluster management; build workload scheduling and data movement systems
Seniority
Senior, hands-on IC