Sr Systems Development Engineer, AWS Hardware Engineering Services, AI UltraServers
Core
Build automation software, diagnostic tooling, and fleet health infrastructure for AWS AI/ML accelerator server fleets to prevent failures and enable zero-touch operations.
Role type
Senior Systems Development Engineer (Hardware/Software/Firmware)
Builds
Automation pipelines, diagnostic frameworks, telemetry correlation engines, and monitoring dashboards for GPU-accelerated compute fleets.
Domain
Cloud infrastructure, AI/ML hardware, Linux systems, and datacenter operations.
Required skills
C++, Python, Golang, Linux kernel debugging, PCIe topology, GPU diagnostics, telemetry analysis, CI/CD pipeline design, root cause analysis, hardware bring-up, device driver development.
Preferred skills
CUDA kernel experience, hardware design validation, leading technical teams, full SDLC ownership, large-scale computing infrastructure. (via careerplan.io/jobs/10569346-sr-systems-development-engineer-aws-hardware-engineering-services-ai-ultraservers-at-amazo)
Technologies
Linux, x86, ARM, PCIe, NVMe, GPU subsystems, CI/CD, telemetry pipelines.
Responsibilities
Design and own automation infrastructure for fleet health; develop test frameworks and diagnostic tooling; build predictive failure detection using telemetry; debug complex system-level issues across compute, GPU, and networking; design scalable test automation for hardware bring-up and qualification.
Seniority
Senior, hands-on IC