Software Engineer- AI/ML, Amazon Neuron Training
Core
Optimize distributed training performance and throughput for large-scale models on AWS Trainium hardware using the Neuron software stack.
Role type
Senior IC distributed systems and ML infrastructure engineer
Builds
High-performance training infrastructure, parallelism strategies, and optimized kernels for Trainium
Domain
Cloud computing, deep learning, high-performance computing (HPC)
Deliverable
production ML models
Required skills
Distributed training optimization, parallelism strategies (data, tensor, pipeline, expert, context), performance profiling, kernel tuning, system architecture design, PyTorch, JAX
Preferred skills
Full software development lifecycle, reduced-precision formats, open source framework contributions
Technologies
PyTorch, JAX, Neuron compiler, Neuron runtime, Trainium
Responsibilities
Lead efforts to optimize distributed training throughput across the Neuron stack; own parallelism strategies for large-scale models; profile workloads to identify bottlenecks (compute, memory, collectives, host overhead); drive fixes across compiler, runtime, and collectives layers; translate performance gaps into requirements for frameworks.
Seniority
Senior, hands-on IC