CareerPlanSign in

Software Engineer- AI/ML, Amazon Neuron Training

Seattle, Washington, United States💼 Full-time💰 $165,200–$165,200🗓 2026-08-13 → 2026-09-26

Core

Optimize distributed training performance and throughput for large-scale models on AWS Trainium hardware using the Neuron software stack.

Role type

Senior IC distributed systems and ML infrastructure engineer

Builds

High-performance training infrastructure, parallelism strategies, and optimized kernels for Trainium

Domain

Cloud computing, deep learning, high-performance computing (HPC)

Deliverable

production ML models

Required skills

Distributed training optimization, parallelism strategies (data, tensor, pipeline, expert, context), performance profiling, kernel tuning, system architecture design, PyTorch, JAX

Preferred skills

Full software development lifecycle, reduced-precision formats, open source framework contributions

Technologies

PyTorch, JAX, Neuron compiler, Neuron runtime, Trainium

Responsibilities

Lead efforts to optimize distributed training throughput across the Neuron stack; own parallelism strategies for large-scale models; profile workloads to identify bottlenecks (compute, memory, collectives, host overhead); drive fixes across compiler, runtime, and collectives layers; translate performance gaps into requirements for frameworks.

Seniority

Senior, hands-on IC

Sourced via amazon · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.