CareerPlanSign in

Sr Systems Development Engineer, AWS AI/ML Servers

Seattle, Washington, United States💼 Full-time🗓 2026-09-11 → 2026-09-25

Core

Build automation, diagnostics, and predictive intelligence for AWS's large-scale AI/ML accelerator server fleet to prevent customer-impacting failures.

Role type

Senior Systems Development Engineer (Hardware/Software/Firmware)

Builds

Automation infrastructure, diagnostic tooling, predictive failure detection systems, and fleet health monitoring dashboards for GPU-accelerated servers.

Domain

Cloud infrastructure, AI/ML hardware, Linux systems, and datacenter operations.

Deliverable

production ML models | product features | dashboards & analysis | infrastructure

Required skills

Systems design, software development, Linux/Unix operations, automation, C++, Python, Golang, hardware debugging, PCIe topology, GPU diagnostics, telemetry pipelines, CI/CD pipelines, device drivers (ARM/x86), root cause analysis, test framework design.

Preferred skills

CUDA kernels, ML/low-level kernels, hardware design and validation, leading technical teams, data center engineering.

Technologies

Linux, C++, Python, Golang, PCIe, GPU, NVMe, ARM, x86, CUDA, CI/CD.

Responsibilities

Build zero-touch automation for fleet health; design diagnostic tooling and test frameworks; develop predictive failure detection using telemetry; create monitoring dashboards and alerting; debug complex system-level issues across compute, GPU, and networking; design scalable test automation for hardware bring-up and qualification.

Seniority

Senior, hands-on IC

Sourced via amazon · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.