Sr. Software Engineer- AI/ML, AWS Neuron
Core
Senior Software Engineer optimizing deep learning and GenAI workloads on AWS custom AI accelerators (Inferentia and Trainium) via the Neuron SDK.
Role type
Senior IC machine-learning systems engineer (AI accelerator optimization)
Builds
Distributed inference support for PyTorch, high-performance kernels, and optimized graph implementations for ML models.
Domain
Cloud computing, Machine Learning, High-Performance Computing, AI Hardware
Deliverable
production ML models
Required skills
Python, C++, System-level programming, Distributed computing architecture, Performance profiling, Low-level optimization, Memory management, Parallel computing, Fusion/Sharding/Tiling/Scheduling
Preferred skills
PyTorch development, Transformer architecture knowledge, Multi-GPU/Multi-node scaling (NCCL), Custom CUDA/Triton kernel development
Technologies
PyTorch, Neuron SDK, AWS Trainium, AWS Inferentia, CUDA, Triton, NCCL
Responsibilities
Design and optimize ML models (GPT, Kimi, Qwen) for custom AI accelerators; Develop high-performance kernels and features for ML operations; Analyze and optimize system-level performance across hardware generations; Implement optimizations like fusion, sharding, tiling, and scheduling; Collaborate with compiler, runtime, and hardware teams; Debug performance issues and resolve bottlenecks.
Seniority
Senior, hands-on IC