Senior Data Center Systems Engineer -- AI Infrastructure & Performance Optimization
Core
Optimizing interdependencies between AI/HPC computing workloads and data center cooling efficiency to maximize performance and energy efficiency.
Role type
Senior IC data center systems engineer (AI infrastructure)
Builds
High-density AI compute clusters with liquid cooling solutions for hyperscalers and OEMs
Domain
Data center infrastructure + AI/HPC hardware
Deliverable
production ML models | infrastructure
Required skills
GPU-based AI workload deployment, liquid cooling hardware commissioning, thermal/power characteristic analysis, digital twin simulation, telemetry data extraction, DCIM/BAS platform integration, stress testing, workload orchestration
Preferred skills
Data center environment experience, NVIDIA/AMD AI hardware architecture knowledge, cross-functional collaboration with thermal engineers
Technologies
NVIDIA H100/H200, AMD MI300, Kubernetes, Slurm, DCIM platforms, BAS platforms, InfiniBand, Ethernet fabrics
Responsibilities
Deploy and manage AI/HPC compute clusters; commission liquid cooling hardware; design workload management strategies; develop digital twin simulations; analyze telemetry data; collaborate on cooling system validation; author technical reports and white papers; represent company at industry conferences; interface with hyperscalers and partners
Seniority
Senior, hands-on IC