Software Engineering Manager, ML Kernel Performance
Core
Design and build high-performance compute kernels for ML operations using the Neuron architecture to optimize deep learning and GenAI workloads on AWS accelerators.
Role type
Senior IC machine-learning kernel engineer (hardware-software boundary)
Builds
High-performance compute kernels and optimization techniques for ML operations on AWS Neuron hardware (Inferentia and Trainium)
Domain
Cloud AI / High-performance computing / Machine Learning
Deliverable
production ML models
Required skills
Kernel-level performance optimization, compiler optimization (fusion, sharding, tiling, scheduling), profiling and bottleneck analysis, system architecture design, multi-tier web services development, cross-functional collaboration
Preferred skills
Experience with Neuron architecture and programming models, deep learning framework integration, startup-like development environment adaptability
Technologies
AWS Neuron, Inferentia, Trainium, PyTorch, AWS Hardware Support, Machine Learning, Web Cloud, AI Architect, Backbone
Responsibilities
Analyze and improve kernel-level performance across multiple generations of Neuron hardware, implement compiler optimizations, work directly with customers to enable and optimize ML models, collaborate across compiler, runtime, framework, and hardware teams, participate in design discussions and code reviews
Seniority
Senior, hands-on IC
