Sr Systems Development Engineer, AWS AI/ML Servers
Core
Build automation, diagnostics, and predictive intelligence for AWS's large-scale AI/ML accelerator server fleet to prevent customer-impacting failures.
Role type
Senior Systems Development Engineer (Hardware/Software/Firmware)
Builds
Automation infrastructure, diagnostic tooling, predictive failure detection systems, and fleet health monitoring dashboards for GPU-accelerated servers.
Domain
Cloud infrastructure, AI/ML hardware, Linux systems, and datacenter operations.
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Systems design, software development, Linux/Unix operations, automation, C++, Python, Golang, hardware debugging, PCIe topology, GPU diagnostics, telemetry pipelines, CI/CD pipelines, device drivers (ARM/x86), root cause analysis, test framework design.
Preferred skills
CUDA kernels, ML/low-level kernels, hardware design and validation, leading technical teams, data center engineering.
Technologies
Linux, C++, Python, Golang, PCIe, GPU, NVMe, ARM, x86, CUDA, CI/CD.
Responsibilities
Build zero-touch automation for fleet health; design diagnostic tooling and test frameworks; develop predictive failure detection using telemetry; create monitoring dashboards and alerting; debug complex system-level issues across compute, GPU, and networking; design scalable test automation for hardware bring-up and qualification.
Seniority
Senior, hands-on IC