Senior Software Development Engineer, AI/ML, AWS Neuron, Model Inference
Core
Architect and implement business-critical features for distributed inference support of LLMs (e.g., Llama, DeepSeek) on Amazon's custom ML accelerators (Trainium and Inferentia), optimizing performance across the stack from PyTorch/JAX to hardware kernels.
Role type
Senior IC Software Development Engineer (AI/ML Systems & Hardware Acceleration)
Builds
High-performance inference kernels, distributed computing architectures, and optimization frameworks for AWS Neuron SDK
Domain
Cloud Infrastructure / Machine Learning / High-Performance Computing / Custom Silicon
Deliverable
production ML models
Required skills
Distributed systems architecture, ML model optimization (latency/throughput), C++ and Python development, system-level programming, memory management, parallel computing, performance profiling, low-level optimization (fusion, sharding, tiling, scheduling)
Preferred skills
Master's degree in CS, full SDLC experience
Technologies
PyTorch, JAX, AWS Trainium, AWS Inferentia, AWS Neuron SDK, C++, Python
Responsibilities
Design and optimize machine learning models and frameworks for custom hardware; Build infrastructure to analyze and onboard diverse model architectures; Implement high-performance kernels leveraging Neuron architecture; Analyze and optimize system-level performance across hardware generations; Conduct detailed performance analysis to resolve bottlenecks; Work directly with customers to enable and optimize ML models on AWS accelerators
Seniority
Senior, hands-on IC with mentorship responsibilities