CareerPlanSign in

Sr. Software Engineer- AI/ML, Amazon Neuron Training

Cupertino, California, United States💼 Full-time💰 $193,300–$261,500🗓 2026-09-22 → 2026-09-25

Core

Lead the effort to build distributed training and post-training support for PyTorch and JAX on AWS Trainium accelerators, enabling large-scale training, post-training, and reinforcement learning workloads.

Role type

Senior IC distributed systems and ML infrastructure engineer

Builds

Distributed training infrastructure, parallelism techniques, high-performance kernels, and optimization strategies for custom ML accelerators

Domain

Cloud computing, machine learning, high-performance computing, distributed systems

Deliverable

production ML models

Required skills

Distributed training, parallelism strategies (data, tensor, pipeline, expert, context), performance profiling, system architecture design, LLM/transformer fundamentals, machine learning frameworks

Preferred skills

RLHF/PPO/DPO experience, production LLM deployment on AI accelerators, performance engineering, open source contributions

Technologies

PyTorch, JAX, AWS Trainium, Neuron compiler, Neuron runtime

Responsibilities

Own parallelism strategies for large-scale models, profile end-to-end workloads to identify bottlenecks, drive fixes across compiler/runtime/collectives layers, translate performance gaps into framework requirements

Seniority

Senior, hands-on IC with leadership responsibilities

Sourced via amazon · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.