Cloud Hardware Development Engineer, Cloud AI/ML server teams
Core
End-to-end owner of accelerator (AI/ML/GPU) server platforms from New Product Introduction (NPI) through fleet health in production, ensuring reliability and operational excellence at scale.
Role type
Senior IC hardware development engineer (cloud AI/ML servers)
Builds
Storage and accelerator server platforms deployed in AWS data center fleets
Domain
Cloud infrastructure, hardware engineering, AI/ML server systems
Deliverable
production ML models | product features | infrastructure
Required skills
server hardware design, NPI lifecycle management, ODM partnership management, functional specification development, design verification planning, root cause analysis, fleet health monitoring, predictive failure detection, system reliability engineering, x86 architecture, high-speed signal integrity
Preferred skills
PCB layout, failure analysis, BIOS/BMC integration, networking hardware, large-scale datacenter operations, storage/compute/GPU platform integration
Technologies
x86, GPU, SSDs, memory, BIOS, BMC, high-speed buses, telemetry systems
Responsibilities
Own end-to-end NPI lifecycle from architecture to launch; design and implement predictive failure detection systems using telemetry; debug complex system failures across hardware, firmware, and physical layers; collaborate with ODMs to manufacture servers at scale; drive zero-touch operations for fleet diagnostics and remediation
Seniority
Senior, hands-on IC