Sr Systems Development Engineer, AWS Hardware Engineering Services, AI UltraServers
Core
Build automation software, diagnostic tooling, and fleet health infrastructure for AWS AI/ML accelerator server fleets to prevent customer-impacting failures.
Role type
Senior Systems Development Engineer (Hardware/Software/Firmware)
Builds
Automation infrastructure, diagnostic tooling, predictive failure detection systems, and monitoring dashboards for accelerated compute fleets.
Domain
Cloud infrastructure, AI/ML hardware, server operations, Linux systems
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
C++, Python, Golang, Linux kernel/driver development, PCIe topology debugging, telemetry pipeline design, CI/CD automation, root cause analysis, ODM collaboration
Preferred skills
Full SDLC leadership, telemetry-based anomaly detection at fleet scale, hardware design partner engagement, large-scale datacenter operations
Technologies
Linux, PCIe, GPU subsystems, BMC/IPMI, ARM/x86, CI/CD pipelines
Responsibilities
Design and develop test frameworks and diagnostic tooling for hardware qualification; Build predictive failure detection using telemetry and sensor data; Debug complex system-level issues across compute, GPU, and networking; Develop scalable test automation for hardware bring-up and regression; Collaborate with ODMs on testability and diagnostic coverage during hardware design.
Seniority
Senior, hands-on IC with architectural responsibilities