Sr. HPC Systems Engineer (IT@JH Research Computing)
Core
Design, build, and maintain advanced high-performance computing (HPC) and AI environments supporting Johns Hopkins University's research mission.
Role type
Senior HPC Systems Engineer
Builds
Multi-node CPU and GPU clusters, high-speed InfiniBand and Ethernet networks, large-scale parallel and object storage
Domain
Academic research computing / High-performance computing infrastructure
Deliverable
production ML models | infrastructure
Required skills
Linux systems administration, cluster management, workload scheduling (Slurm), distributed storage administration, high-speed interconnects (Infiniband), automation scripting (Bash, Python), performance monitoring and tuning, security hardening
Preferred skills
Configuration management (Ansible, Puppet, Salt), containerization and orchestration (Singularity, Docker, Kubernetes), GPU acceleration, research software environments (CUDA, MPI, AI/ML frameworks), capacity planning, hardware lifecycle management
Technologies
Slurm, Infiniband, Ethernet, GPFS, Lustre, WekaFS, Ceph, MinIO, Prometheus, Grafana, ELK, Ansible, Puppet, Salt, Singularity, Docker, Kubernetes, CUDA, MPI
Responsibilities
Support and administer production systems used by researchers; Provide technical leadership for system configuration and implementation; Research and recommend new functionality for HPC management tools; Architect, operate, and debug large-scale HPC network and storage infrastructure; Analyze server monitoring results to improve performance and utilization; Propose and enforce security policies; Deploy and maintain large-scale Linux-based HPC clusters; Implement and optimize workload schedulers; Administer distributed storage systems; Maintain high-speed fabric and network infrastructure; Develop and maintain automation and monitoring frameworks; Participate in capacity planning and hardware lifecycle management
Seniority
Senior, hands-on IC