Sr. Software Engineer- AI/ML, Amazon Neuron Training
Core
Lead the effort to build distributed training and post-training support for PyTorch and JAX on AWS Trainium accelerators, enabling large-scale training, post-training, and reinforcement learning workloads.
Role type
Senior IC distributed systems and ML infrastructure engineer
Builds
Distributed training infrastructure, parallelism techniques, high-performance kernels, and optimization strategies for custom ML accelerators
Domain
Cloud computing, machine learning, high-performance computing, distributed systems
Deliverable
production ML models
Required skills
Distributed training, parallelism strategies (data, tensor, pipeline, expert, context), performance profiling, system architecture design, LLM/transformer fundamentals, machine learning frameworks
Preferred skills
RLHF/PPO/DPO experience, production LLM deployment on AI accelerators, performance engineering, open source contributions
Technologies
PyTorch, JAX, AWS Trainium, Neuron compiler, Neuron runtime
Responsibilities
Own parallelism strategies for large-scale models, profile end-to-end workloads to identify bottlenecks, drive fixes across compiler/runtime/collectives layers, translate performance gaps into framework requirements
Seniority
Senior, hands-on IC with leadership responsibilities