Software Engineering Manager, ML Kernel Performance, AWS Neuron, Annapurna Labs
Core
Architect and implement high-performance compute kernels for ML operations on AWS custom accelerators (Inferentia/Trainium), optimizing deep learning and GenAI workloads at the hardware-software boundary.
Role type
Senior IC machine-learning kernel engineer (hardware acceleration)
Builds
High-performance ML inference and training kernels for AWS Neuron SDK
Domain
Machine learning, high-performance computing, custom hardware acceleration
Deliverable
production ML models
Required skills
Low-level optimization, system architecture, ML model acceleration, compiler optimization (fusion, sharding, tiling, scheduling), performance profiling, hardware-software integration
Preferred skills
Customer enablement for ML models, cross-functional collaboration, startup-like development environment adaptability
Technologies
AWS Neuron, Inferentia, Trainium, PyTorch, Neuron Compiler, Neuron Runtime
Responsibilities
Design and implement high-performance compute kernels for ML operations; Analyze and optimize kernel-level performance across Neuron hardware generations; Conduct detailed performance analysis using profiling tools; Implement compiler optimizations; Work directly with customers to enable and optimize ML models; Collaborate across teams to develop innovative kernel optimization techniques
Seniority
Senior, hands-on IC