Staff ML Engineer – AWS Trainium & SageMaker
Core
Operate and train large-scale AI models on Amazon SageMaker using AWS Trainium custom silicon, bridging hardware-level compiler behavior with PyTorch production pipelines.
Role type
Staff ML Engineer (AWS Trainium & SageMaker)
Builds
Production training pipelines and optimized model training workloads for enterprise clients
Domain
Cloud AI Infrastructure / Custom Silicon Accelerators
Deliverable
production ML models
Required skills
PyTorch (distributed/multi-device), Amazon SageMaker, Python, hardware-aware debugging, cost-aware pipeline design
Preferred skills
AWS Trainium/Inferentia (Neuron SDK), NeuronCore architecture, compiler behavior understanding
Technologies
Amazon SageMaker, AWS Trainium, PyTorch, Neuron SDK
Responsibilities
Write and optimize PyTorch training code for Trainium hardware, diagnose hardware-specific training issues, tune distributed training for throughput and cost, translate training requests into end-to-end production pipelines, collaborate with client teams on production workloads
Seniority
Staff, hands-on IC with client-facing delivery
