Systems Development Engineer, AWS Generative AI & ML Servers
Core
Build automation software, diagnostic tooling, and fleet health infrastructure for AWS accelerated (AI/ML) server platforms to enable zero-touch operations.
Role type
Senior Systems Development Engineer (Server Hardware & Automation)
Builds
Automation infrastructure, predictive failure detection systems, monitoring tools, and diagnostic tooling for AI/ML server fleets.
Domain
Cloud Infrastructure / AI/ML Server Hardware / Systems Engineering
Deliverable
production ML models | infrastructure
Required skills
Linux OS internals, C/C++/Python/Java/Golang, system architecture & design, root cause analysis, device driver development, CI/CD pipeline management, telemetry & log correlation, hardware diagnostics (PCIe, GPU, NVMe), fleet health metrics definition.
Preferred skills
CUDA kernels, BMC/IPMI familiarity, zero-touch/self-healing automation concepts, ODM collaboration, large-scale datacenter operations, predictive maintenance procedures.
Technologies
Linux (x86/ARM), Python, Ruby, Java, C/C++, Golang, PCIe, NVMe, GPU subsystems, CI/CD tools.
Responsibilities
Build automation for fleet health and predictive failure detection; debug complex system-level issues across compute/GPU/networking; develop device drivers and diagnostic tooling; define and track fleet health metrics; collaborate with HWEng teams and ODMs on server design and testability.
Seniority
Senior, hands-on IC