Supercomputing Intern
Core
Design, development, and deployment of ML system software for operating rack-scale supercomputing systems, focusing on network performance, telemetry pipelines, and system health analysis.
Role type
Supercomputing Intern (ML System Software)
Builds
Rack-scale ML systems, telemetry processing pipelines, and software frameworks for data center scale inference workloads.
Domain
Hardware/Software Co-design, Supercomputing, Frontier AI Infrastructure
Deliverable
production ML models | infrastructure
Required skills
C/C++ or Rust, Python, data structures and algorithms, low-level software engineering, hardware/software co-design
Preferred skills
Linux internals, kernel development, driver debugging, hardware diagnostics, server virtualization, CI/CD pipelines, embedded development
Technologies
C, C++, Rust, Python, Linux
Responsibilities
Design and develop ML system software for rack-scale systems, create and process telemetry pipelines, analyze system-level health and performance, deploy and provision software frameworks, validate hardware.
Seniority
Intern
