Systems Development Engineer, GPU & AI Accelerator Servers, AWS Hardware Engineering
Core
Develop automation, analyze hardware telemetry, and build tooling to ensure the health and availability of server platforms at worldwide fleet scale for Generative AI customers.
Role type
Senior Systems Development Engineer (Hardware/Software)
Builds
Fleet-wide data pipelines, operational dashboards, and automation for hardware test, firmware qualification, and capacity recovery workflows.
Domain
Cloud infrastructure, AI accelerator servers, hardware telemetry, and fleet operations.
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Systems engineering fundamentals, hardware failure analysis, predictive failure detection, Linux on ARM and x86 debugging, automation development, fleet telemetry analysis, component lifecycle tracking, cross-team collaboration.
Preferred skills
Agile/Scrum methodology, PowerShell scripting.
Technologies
Python, Java, C++, C#, Golang, PowerShell, Ruby, Linux, ARM, x86, PCIe, NVMe, GPU subsystems.
Responsibilities
Analyze hardware failure patterns using fleet telemetry and system event logs; build and maintain operational dashboards for platform fleet health; develop diagnostic tools for Linux on ARM and x86 architectures; debug and resolve Linux boot and runtime issues across processor architectures; collaborate with software, hardware, and vendor teams to validate new compute solutions.
Seniority
Senior, hands-on IC