Infrastructure Engineer
Core
Ensuring reliability and stability of GPU hardware platforms through hands-on diagnosis, repair, and automation development.
Role type
Senior IC infrastructure engineer (hardware operations)
Builds
Hardware diagnostics, provisioning, and repair tooling for AI compute fleets
Domain
AI infrastructure / Data center hardware operations
Deliverable
production ML models | infrastructure
Required skills
Linux internals, server hardware troubleshooting, hardware and networking fundamentals, automation scripting, system log analysis, on-call incident response
Preferred skills
Large-scale GPU operations experience, proficiency in Python or Go
Technologies
Linux, NVIDIA GPUs, AMD GPUs, BMC Redfish APIs
Responsibilities
Investigate and troubleshoot hardware faults in GPU platforms, automate routine fleet management processes, validate and test new generation AI hardware, collaborate with operations teams on hardware remediation, build documentation and tooling for knowledge sharing, participate in on-call rotation for follow-the-sun coverage
Seniority
Mid-to-Senior, hands-on IC