CareerPlanSign in

Senior Software Development Engineer, AI/ML, AWS Neuron, Model Inference

Cupertino, California, United States💼 Full-time🗓 2026-08-18 → 2026-09-25

Core

Architect and implement business-critical features for distributed inference support of LLMs (e.g., Llama, DeepSeek) on Amazon's custom ML accelerators (Trainium and Inferentia), optimizing performance across the stack from PyTorch/JAX to hardware kernels.

Role type

Senior IC Software Development Engineer (AI/ML Systems & Hardware Acceleration)

Builds

High-performance inference kernels, distributed computing architectures, and optimization frameworks for AWS Neuron SDK

Domain

Cloud Infrastructure / Machine Learning / High-Performance Computing / Custom Silicon

Deliverable

production ML models

Required skills

Distributed systems architecture, ML model optimization (latency/throughput), C++ and Python development, system-level programming, memory management, parallel computing, performance profiling, low-level optimization (fusion, sharding, tiling, scheduling)

Preferred skills

Master's degree in CS, full SDLC experience

Technologies

PyTorch, JAX, AWS Trainium, AWS Inferentia, AWS Neuron SDK, C++, Python

Responsibilities

Design and optimize machine learning models and frameworks for custom hardware; Build infrastructure to analyze and onboard diverse model architectures; Implement high-performance kernels leveraging Neuron architecture; Analyze and optimize system-level performance across hardware generations; Conduct detailed performance analysis to resolve bottlenecks; Work directly with customers to enable and optimize ML models on AWS accelerators

Seniority

Senior, hands-on IC with mentorship responsibilities

Sourced via amazon · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.