Senior ML Performance Engineer, Annapurna Labs
Core
Building the AWS Neuron software stack to optimize performance for Generative AI and advanced ML workloads on custom ML accelerators (Inferentia and Trainium).
Role type
Senior IC machine learning performance engineer (systems & hardware)
Builds
High-performance kernels, Neuron SDK enhancements, and optimized ML software stack for cloud inference and training
Domain
Cloud infrastructure, ML systems, hardware-software co-design
Deliverable
production ML models
Required skills
C++, Python, PyTorch, TensorFlow, JAX, performance engineering for accelerated computing, kernel optimization, instruction scheduling, memory management, parallelism, compiler enhancements
Preferred skills
Compiler optimization, hardware-software co-design, LLM/Vision deep-learning model optimization
Technologies
AWS Neuron, Inferentia, Trainium, PyTorch, TensorFlow, JAX
Responsibilities
Optimize system performance across the ML software stack, analyze high-performance ML workloads on Annapurna hardware, develop high-performance kernels for critical ML operations, enhance the Neuron SDK, collaborate across Compiler, Frameworks, and Hardware teams
Seniority
Senior, hands-on IC