Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training
Core
Architect and implement business-critical features for AWS Neuron to accelerate deep learning and GenAI workloads on AWS Trainium, focusing on distributed training performance, model enablement, and high-performance kernel optimization.
Role type
Senior IC machine-learning engineer (distributed training & HPC)
Builds
Infrastructure for large-scale training on custom ML accelerators, including parallelism strategies, numerics techniques, and high-performance kernels.
Domain
Cloud infrastructure, Machine Learning, High-Performance Computing, Distributed Systems (via careerplan.io/jobs/10497224-sr-software-engineer-aiml-aws-neuron-distributed-training-at-amazon)
Deliverable
production ML models | infrastructure
Required skills
Distributed training optimization, PyTorch, JAX, Neuron compiler/runtime, parallelism strategies (data/tensor/pipeline/expert/context), reduced-precision formats, performance profiling, system architecture design, mentorship
Preferred skills
Computer architecture, PyTorch/JAX/Tensorflow expertise, End-to-end model training experience
Technologies
AWS Trainium, PyTorch, JAX, Neuron SDK, AWS Neuron compiler, AWS Neuron runtime
Responsibilities
Optimize distributed training throughput and time to convergence, bring up novel model architectures on Trainium, profile workloads to identify bottlenecks, translate performance gaps into future architecture requirements, contribute upstream to open source frameworks
Seniority
Senior, hands-on IC with mentorship responsibilities