Machine Learning Performance Engineer, Annapurna Labs
Core
Building AWS Neuron performance engineering team to profile and optimize deep learning workloads on Inferentia and Trainium chips.
Role type
Senior IC machine learning performance engineer (hardware-software boundary)
Builds
High-performance compute kernels, Neuron SDK enhancements, and optimization strategies for ML workloads
Domain
AI infrastructure, custom ML accelerators, deep learning frameworks
Deliverable
production ML models
Required skills
Python, C++, computer architecture, parallel computing, PyTorch, TensorFlow, JAX
Preferred skills
LLM/Vision model optimization, kernel writing, compiler optimization, hardware-software co-design
Technologies
Inferentia, Trainium, CUDA, Triton, CUTLASS, Pallas, Mojo, SIMD, MPI
Responsibilities
Design and implement high-performance compute kernels for ML operations; Profile ML workloads end-to-end to identify bottlenecks; Enhance programming model and tooling for kernel and model developers; Identify optimization opportunities across Neuron software stack; Document software designs and performance findings
Seniority
Senior, hands-on IC