Data Center Engineer, HPC and AI
Core
Building, maintaining, and operating large-scale HPC and AI supercomputer clusters and data center infrastructure.
Role type
Senior IC Data Center Engineer (HPC/AI Infrastructure)
Builds
Large-scale compute and Deep Learning hardware and software platforms
Domain
High-Performance Computing (HPC) and Artificial Intelligence Infrastructure
Deliverable
infrastructure
Required skills
Linux troubleshooting, rack stacking, cable management, power and cooling optimization, network troubleshooting, OS installation, data center operations
Preferred skills
Bash/Python scripting, configuration management (Ansible, Puppet), CI/CD and job schedulers (Jenkins, SLURM), virtualization (KVM, VMware, Hyper-V), L2/L3 network protocols
Technologies
Linux, Windows, DHCP, DNS, NIS, AD, KVM, VMware, Hyper-V, Jenkins, SLURM, Ansible, Puppet
Responsibilities
Plan and build complex cluster and supercomputers, manage rack stacking and cabling, optimize data center power and cooling efficiency, perform daily operations and support, install infrastructure solutions (Cloud, VMs, Storage, Network, HPC, AI), troubleshoot network and bare metal systems, support R&D activities
Seniority
Senior, hands-on IC