Support Engineer - AI Server Systems
Core
Maintaining and troubleshooting high-availability AI server infrastructure, including GPU clusters, storage, and networking equipment.
Role type
Support Engineer (AI Infrastructure)
Builds
AI server systems and related infrastructure
Domain
AI hardware / High-performance computing
Deliverable
physical/clinical work
Required skills
Hardware troubleshooting (power, memory, storage, PCIe, GPU), Linux server administration, Network fundamentals (L2/L3, TCP/IP, DHCP, IPMI), On-site support execution, Incident reporting, Firmware/BIOS/Driver updates, Inventory management, Installation/migration support
Preferred skills
NVIDIA GPU server experience (DGX, HGX), Data center operations, GPU deep learning basics, Linux shell scripting, English technical negotiation
Technologies
Linux (Ubuntu, RHEL, CentOS), IPMItool, smartctl, nvidia-smi, Ethernet, InfiniBand, NVLink, PCIe switches
Responsibilities
Perform primary troubleshooting and on-site repairs for server failures including component replacement; Monitor system status and analyze logs via NOC and remote tools; Create incident reports and escalate to engineering teams; Execute firmware, BIOS, and driver updates; Manage spare parts inventory and coordinate shipments; Conduct customer site inspections and preventive maintenance; Oversee installation and relocation operations
Seniority
Individual Contributor