Assoc. Dir. DDIT IES - HPC Infrastructure Sr. Engineer
Core
Operates and manages High-Performance Computing (HPC) infrastructure, ensuring service deployments, operational health, and lifecycle management for complex computing solutions.
Role type
Senior HPC Infrastructure Engineer
Builds
HPC clusters and infrastructure for scientific and research workloads
Domain
High-Performance Computing / Scientific Computing
Deliverable
production ML models | infrastructure
Required skills
HPC scheduling (Grid Engine, Slurm, LSF), cluster management (Bright Cluster Manager, xCat, OpenHPC), Linux system administration, Bash scripting, Python, build systems (Make, CMake), parallel file systems (NFS, GPFS, Lustre), InfiniBand networking, Unix performance tuning, DevOps tools (Ansible, Chef, CI/CD)
Preferred skills
AI/ML workload infrastructure setup, HPC in pharmaceutical/bioinformatics environments
Technologies
AWS, GCP, Azure, InfiniBand, Bash, Python, Make, CMake, Grid Engine, Slurm, LSF, Bright Cluster Manager, xCat, OpenHPC, NFS, GPFS, Gluster, BeeGFS, Lustre, Ansible, Chef
Responsibilities
Execute and manage HPC service deployments and lifecycle; design innovative solutions to optimize complex infrastructures; act as escalation point for critical incidents; track suppliers and partners for project delivery; contribute to service and platform strategy development
Seniority
Senior, hands-on IC