CareerPlanSign in

Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training

Cupertino, California, United States💼 Full-time🗓 2026-08-10 → 2026-10-02

Core

Architect and implement business-critical features for AWS Neuron to accelerate deep learning and GenAI workloads on AWS Trainium, focusing on distributed training performance, model enablement, and high-performance kernel optimization.

Role type

Senior IC machine-learning engineer (distributed training & HPC)

Builds

Infrastructure for large-scale training on custom ML accelerators, including parallelism strategies, numerics techniques, and high-performance kernels.

Domain

Cloud infrastructure, Machine Learning, High-Performance Computing, Distributed Systems (via careerplan.io/jobs/10497224-sr-software-engineer-aiml-aws-neuron-distributed-training-at-amazon)

Deliverable

production ML models | infrastructure

Required skills

Distributed training optimization, PyTorch, JAX, Neuron compiler/runtime, parallelism strategies (data/tensor/pipeline/expert/context), reduced-precision formats, performance profiling, system architecture design, mentorship

Preferred skills

Computer architecture, PyTorch/JAX/Tensorflow expertise, End-to-end model training experience

Technologies

AWS Trainium, PyTorch, JAX, Neuron SDK, AWS Neuron compiler, AWS Neuron runtime

Responsibilities

Optimize distributed training throughput and time to convergence, bring up novel model architectures on Trainium, profile workloads to identify bottlenecks, translate performance gaps into future architecture requirements, contribute upstream to open source frameworks

Seniority

Senior, hands-on IC with mentorship responsibilities