Training, Process Management Engineer
Core
Design, build, and maintain high-performance distributed systems to orchestrate and monitor machine learning workloads on supercomputers, ensuring reliability and observability for research experiments and frontier-scale training runs.
Role type
Senior IC distributed systems engineer (process management)
Builds
Distributed OS and runtime components for launching, coordinating, and supervising massive training workloads
Domain
AI infrastructure / High-performance computing / Distributed systems
Deliverable
production ML models
Required skills
Rust, Python, Linux, asynchronous programming, distributed systems design, performance profiling, systems-level debugging, fault tolerance design
Preferred skills
C++, memory profiling, concurrent systems development, large-scale cluster debugging
Technologies
Rust, Python, Linux
Responsibilities
Design and build software to orchestrate and monitor ML workloads on supercomputers; Profile and optimize the software stack for frontier-scale computation; Improve reliability, observability, and fault tolerance for long-running jobs; Debug complex distributed systems issues across large clusters; Respond to evolving needs of ML systems to enable researchers
Seniority
Senior, hands-on IC