Sr. System Development Engineer, High-Performance Accelerator Servers for AI/ML
Core
Design, deliver, and operate next-generation infrastructure for AI training and inference, focusing on high-performance accelerator servers and HPC workloads.
Role type
Senior System Development Engineer (Hardware/Software Systems)
Builds
AWS Accelerated server solutions for the AWS Cloud
Domain
Cloud computing, AI/ML infrastructure, High-Performance Computing (HPC)
Deliverable
production ML models | infrastructure
Required skills
Linux/Unix systems development, C++/C#/Java/Python/Golang/PowerShell/Ruby programming, x86 architecture, systems design and architecture, automation, process improvement, systems debugging, server system testability and reliability
Preferred skills
Full software/hardware/networks development life cycle knowledge, complex software or computing infrastructure delivery, debugging and validating complex AI/ML and Cloud Computing servers
Technologies
Linux/Unix, x86 architecture, C++, C#, Java, Python, Golang, PowerShell, Ruby, Bash, Perl
Responsibilities
Decompose complex server system testability, reliability, and diagnosis problems; lead design, build, and deployment of complex and performant software solutions; collaborate across hardware, software, and network engineering teams; drive high quality and reliability into future server designs
Seniority
Senior, hands-on IC with leadership responsibilities