ML Kernel Performance Engineer
Core
Design and implement high-performance compute kernels for machine-learning operations on AWS Neuron architecture.
Role type
Senior IC ML Kernel Performance Engineer
Builds
Optimized ML inference/training workloads on AWS accelerators
Domain
Cloud computing / Machine Learning / Hardware acceleration
Deliverable
production ML models
Required skills
kernel-level performance optimization, profiling, compiler optimization (fusion, sharding, tiling, scheduling), system architecture, software development
Preferred skills
full software development life-cycle experience, coding standards, code reviews, source control management, build processes, testing, operations
Technologies
Neuron architecture, AWS accelerators
Responsibilities
Analyze and optimize kernel-level performance across Neuron hardware generations, implement compiler optimizations, work directly with customers to enable and optimize ML models, collaborate with compiler/runtime/framework/hardware teams, create metrics and implement automation, publish research and mentor experienced engineers
Seniority
Senior, hands-on IC with mentorship responsibilities