ML Accelerator Performance Validation Engineer, Post Silicon Validation
Core
Quantify and qualify the performance of AWS's custom ML training chips against architectural targets by bridging silicon capabilities with real-world ML workload demands.
Role type
ML accelerator performance validation engineer
Builds
Next-generation AI/ML hardware for AWS training and inference infrastructure
Domain
Cloud computing, AI/ML hardware, silicon validation
Deliverable
production ML models
Required skills
Machine Learning and Large Language Model fundamentals, hardware performance counters and profiling tools, computer architecture fundamentals, statistical methods and regression analysis, Python, C++, Java
Preferred skills
CUDA kernels development, LLM deployment on GPUs/Neurons/TPUs, collective communications (AllReduce, AllGather), HBM/PCIe/DMA bandwidth characterization
Technologies
PyTorch, JAX, CUDA, GPUs, Neuron, TPU
Responsibilities
Design and execute performance benchmarks from micro-architectures to full model training, measure and analyze compute throughput and memory bandwidth, profile real ML workloads on silicon, build automated performance regression dashboards, correlate silicon measurements against RTL simulation
Seniority
Mid-level, hands-on IC