Training, Process Management Engineer
Core
Design, build, and maintain software to orchestrate and monitor machine learning workloads on large supercomputers, ensuring reliability and performance for distributed training runs.
Role type
Senior IC systems engineer (distributed runtime)
Builds
Distributed OS and runtime components for launching, coordinating, and supervising massive training workloads
Domain
AI infrastructure / High-performance computing
Deliverable
production ML models
Required skills
Rust, Python, Linux, asynchronous programming, distributed systems design, performance profiling, systems-level debugging, memory management, fault tolerance design
Preferred skills
Experience developing (not just operating) distributed systems, strong design judgment for ambiguous problems, high-ownership mindset
Technologies
Rust, Python, Linux
Responsibilities
Design and build software to orchestrate and monitor ML workloads on supercomputers; Profile and optimize the software stack for frontier-scale computation; Improve reliability, observability, and fault tolerance for long-running jobs; Debug complex distributed systems issues across large clusters; Respond to evolving needs of ML systems to enable researchers
Seniority
Senior, hands-on IC