L3 Support Engineer
Core
Senior technical expert leading deep investigations, root cause analysis, and permanent fixes for datacenter servers, firmware, and Linux diagnostics across Europe and the US.
Role type
Senior L3 Support Engineer (Datacenter Infrastructure)
Builds
Scalable solutions for server stability, firmware lifecycle management, and platform readiness for AI cloud infrastructure.
Domain
Datacenter hardware, server firmware (BIOS/BMC), Linux systems, and GPU infrastructure.
Deliverable
Production infrastructure stability and scalable troubleshooting playbooks.
Required skills
Deep Linux troubleshooting, datacenter server diagnostics, hardware/firmware interaction analysis, incident response, vendor escalation management, runbook creation, travel for on-site troubleshooting.
Preferred skills
GPU server platform familiarity (NVIDIA tooling), IPMI/Redfish workflows, firmware lifecycle management, Bash/Python scripting, OCP-based platforms, enterprise bare metal SLA support.
Technologies
Linux, BIOS/BMC, NVIDIA tools (nvidia-smi, dcgmi), IPMI, Redfish, Bash, Python.
Responsibilities
Lead root cause analysis for GPU failures, firmware issues, and Linux-level faults; detect recurring patterns across sites; drive evidence-based escalations with ODM and R&D; support firmware update validation and rollout; create scalable troubleshooting guides and error catalogs; travel to datacenters for complex troubleshooting.
Seniority
Senior, hands-on IC with cross-site pattern detection and mentorship of L1/L2 teams.