AI Hardware Systems Engineer, Annapurna Labs, Trainium Machine Learning Fleet Operations
Core
Debug emergent problems in GPU and server hardware, run large-scale experiments, and develop automation software to scale ML fleet operations.
Role type
Senior IC systems engineer (ML fleet operations)
Builds
Autonomous software for monitoring, optimizing, and remediating machine learning hardware fleets
Domain
Cloud infrastructure / Machine Learning hardware / Systems engineering
Deliverable
production ML models | infrastructure
Required skills
Python, Bash, C++, Linux/Unix, systems engineering fundamentals, debugging, data analysis, automation development
Preferred skills
Hardware design and validation, post-silicon validation
Technologies
Python, Bash, Linux, C++, Golang
Responsibilities
Review dashboards to identify trends and triage emergent issues, partner with teams to debug hardware/software issues, implement system-level testing, develop maintainable and reusable software, direct new automations and data infrastructure
Seniority
Senior, hands-on IC