Datacentre Operations Engineer
Core
Hands-on operations engineer for high-density, air-cooled AI/HPC compute infrastructure, focusing on hardware deployment, break/fix, cooling management, and 24/7 site reliability.
Role type
Senior IC Datacentre Operations Engineer (HPC/AI Infrastructure)
Builds
Next-generation NVIDIA B300-class GPU compute platforms and high-density air-cooled systems for AI and HPC workloads
Domain
Data Centre Operations / High-Performance Computing / AI Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Hardware troubleshooting, break/fix procedures, high-density air cooling operations, capacity management, asset lifecycle management, incident response, firmware updates, busbar power distribution, thermal monitoring
Preferred skills
Experience with NVIDIA GPU platforms, HPC environments, ITSM frameworks (Incident/Change Management), automation, cross-training in SRE practices
Technologies
NVIDIA B300-class chassis, CRAC/CRAH units, busbar power distribution, PDU infrastructure, monitoring and observability platforms
Responsibilities
Diagnose and resolve hardware/network issues to maximize uptime; execute break/fix procedures including GPU module exchange and component replacement; operate and maintain high-density air cooling infrastructure; manage site-level capacity (power, space, cooling); handle RMAs and support requests within SLOs; contribute to incident postmortems and root cause analysis; maintain SOPs for safe working around high-density systems
Seniority
Senior, hands-on IC