Lead Systems Engineer HPC
Core
Design, maintain, and improve advanced HPC and AI cluster infrastructure, including high-performance interconnects, schedulers, and configuration management for research systems.
Role type
Lead Systems Engineer HPC
Builds
Advanced HPC and AI cluster infrastructure, data-transfer pathways, and networking solutions for AI-driven computing workloads.
Domain
Academic research computing, High-Performance Computing (HPC), Artificial Intelligence (AI)
Deliverable
production ML models | infrastructure
Required skills
Linux system administration, Bash scripting, Python scripting, Perl scripting, HPC networking administration, SLURM job scheduling, cluster software management, complex systems troubleshooting, technical strategy definition, team mentorship
Preferred skills
Academic and research environment experience, AI-driven research support, Globus data transfer tools, parallel file systems, unstructured data support
Technologies
AI, Bash, Hardware Support, Linux, Perl, Python, Security
Responsibilities
Design, maintain, and troubleshoot advanced HPC and AI cluster infrastructure; develop data-transfer pathways and networking solutions; establish best practices for cluster administration; create user and technical documentation; expand monitoring infrastructure; plan and carry out scheduled maintenance; define and guide technical strategy for AI and data-intensive HPC; mentor systems specialists and analysts; advise senior leadership on investments and risks; monitor clusters, networks, and storage systems for irregularities; analyze and fix problems in Linux and HPC/AI environments.
Seniority
Senior, hands-on IC with leadership and mentorship responsibilities