Software Engineer II - AI/ML, Neuron Inference
Core
Design, develop, and optimize machine learning models and frameworks for deployment on custom ML hardware accelerators (AWS Trainium), focusing on distributed inference and high-performance kernels.
Role type
Senior IC software engineer (AI/ML inference optimization)
Builds
Distributed inference solutions for PyTorch and large-scale LLMs on AWS Trainium
Domain
Cloud infrastructure, Machine Learning, High-Performance Computing, Custom Hardware Acceleration
Deliverable
production ML models
Required skills
Python, C++, distributed computing, system-level programming, ML model optimization, performance profiling, hardware architecture knowledge, parallel computing, memory management
Preferred skills
CUDA kernel development, Triton syntax, vLLM, SGLang, TensorRT, JIT compilation, computer architecture
Technologies
PyTorch, AWS Trainium, Neuron SDK, CUTLASS, FlashInfer
Responsibilities
Design and implement high-performance kernels for ML operations; optimize system-level performance across Neuron hardware generations; build infrastructure to onboard diverse model architectures; conduct performance analysis and resolve bottlenecks; implement optimizations like fusion, sharding, and tiling; collaborate with compiler and runtime teams; work directly with customers to enable model performance.
Seniority
Senior, hands-on IC