HPC AI Systems Administrator Lead
Core
Senior System Administrator providing advanced system administration and lab operations support for hardware, network, and software environments used by HPE HPC & AI Performance Engineering teams.
Role type
Senior System Administrator (HPC & AI Lab Operations)
Builds
Internal product development, performance engineering, ISV validation, and customer-facing sales and benchmarking activities
Domain
High Performance Computing (HPC) and Artificial Intelligence (AI) Infrastructure
Deliverable
infrastructure
Required skills
Linux system administration, HPC cluster provisioning, virtual server administration, high-performance storage management, hardware diagnostics, job scheduling, capacity planning
Preferred skills
Network administration, advanced lab system administration
Technologies
Linux, Lustre, virtualization platforms, job scheduling tools
Responsibilities
Image, configure, and upgrade servers with Linux operating systems; Configure and manage multiple root slots for HPC cluster provisioning; Provide design guidance for virtualized lab infrastructure and high-performance storage solutions; Collaborate with AI benchmarking and R&D teams to design and operate lab environments; Oversee lab transitions, facility moves, and infrastructure refresh activities; Serve as a technical mentor to junior system administrators and lab staff.
Seniority
Senior, hands-on IC with mentorship responsibilities