Senior ML Kernel Performance Engineer
Core
Architect and implement high-performance compute kernels for ML operations on Amazon's custom ML accelerators (Inferentia and Trainium), optimizing workloads at the hardware-software boundary.
Role type
Senior IC machine-learning kernel performance engineer
Builds
High-performance compute kernels, compiler optimizations, and runtime solutions for deep learning and GenAI workloads
Domain
Machine learning, high-performance computing, custom hardware acceleration
Deliverable
production ML models
Required skills
Low-level optimization, system architecture, ML model acceleration, profiling and bottleneck resolution, compiler optimization (fusion, sharding, tiling, scheduling), cross-functional collaboration
Preferred skills
Accelerator architecture expertise (GPUs, CPUs, FPGAs, custom), GPU kernel optimization (CUDA, NKI, Triton, OpenCL, SYCL, ROCm), LLVM/MLIR backend development, parallel programming, GPU memory hierarchy optimization
Technologies
Neuron architecture, PyTorch, TensorFlow, CUDA, LLVM, MLIR
Responsibilities
Design and implement high-performance compute kernels for ML operations; Analyze and optimize kernel-level performance across Neuron hardware generations; Conduct detailed performance analysis using profiling tools; Implement compiler optimizations; Work directly with customers to enable and optimize ML models; Collaborate across teams to develop innovative kernel optimization techniques
Seniority
Senior, hands-on IC with mentorship responsibilities