Software Development Engineer, ML Acceleration, Trainium AI Systems, Annapurna Labs
Core
Build data pipelines, metrics, alarms, and dashboards to analyze diagnostic test results from tens of thousands of ML accelerator servers, distinguishing genuine hardware faults from test noise to gate production readiness.
Role type
Senior IC software engineer (data engineering & analytics)
Builds
Production data pipelines, fleet-wide monitoring dashboards, and anomaly detection systems for hardware diagnostic telemetry (via careerplan.io/jobs/10570092-software-development-engineer-ml-acceleration-trainium-ai-systems-annapurna-labs-at-amazon)
Domain
Semiconductor hardware validation / Machine Learning infrastructure / Data engineering
Required skills
Python, SQL, large-scale data pipeline design, workflow orchestration, anomaly detection, time-series analysis, data warehouse management
Preferred skills
Telemetry/diagnostic data experience, server hardware familiarity, Grafana/QuickSight dashboarding
Technologies
Amazon Redshift, Apache Spark, Grafana, QuickSight
Responsibilities
Design and operate pipelines ingesting diagnostic and repair-ticket data from the Trainium/Inferentia fleet; Build metrics and alarms to detect test regressions and failure-rate shifts; Create visualizations for fleet health, failure signatures, and yield metrics; Analyze fleet data to separate hardware faults from software defects and test noise; Define evidence standards for moving diagnostics from observation to production blocking; Improve data quality and pipeline reliability for downstream consumers.
Seniority
Senior, hands-on IC