CareerPlanSign in

Software Engineer, ML Infrastructure

San Francisco💼 Full-time🗓 2026-01-27 → 2026-09-26

Core

Building large-scale compute, storage, and software infrastructure to support training the world's best agentic coding model.

Role type

Senior IC ML Infrastructure Engineer

Builds

High-performance GPU infrastructure, training frameworks, and workload scheduling systems for agentic coding models

Domain

AI/ML Infrastructure, Distributed Systems, Cloud & Bare Metal

Deliverable

infrastructure

Required skills

Python, TypeScript, Rust, Golang, distributed storage, Linux systems, Kubernetes, infrastructure-as-code, workload scheduling, data movement systems

Preferred skills

Nvidia GPU operations (Infiniband/RoCE), Ray, Slurm

Technologies

Kubernetes, Linux, Python, TypeScript, Rust, Golang, Ray, Slurm

Responsibilities

Collaborate with ML researchers to improve training throughput and reliability; plan and build cutting-edge GPU infrastructure with OEMs and cloud providers; improve compute density and scalability for large RL workloads; create software to automate GPU cluster management; build workload scheduling and data movement systems

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.