CareerPlanSign in

Staff ML Engineer – AWS Trainium & SageMaker

Canada - Remote💼 Full-time🗓 2026-09-16 → 2026-09-26

Core

Operate and train large-scale AI models on Amazon SageMaker using AWS Trainium custom silicon, bridging hardware-level compiler behavior with PyTorch production pipelines.

Role type

Staff ML Engineer (AWS Trainium & SageMaker)

Builds

Production training pipelines and optimized model training workloads for enterprise clients

Domain

Cloud AI Infrastructure / Custom Silicon Accelerators

Deliverable

production ML models

Required skills

PyTorch (distributed/multi-device), Amazon SageMaker, Python, hardware-aware debugging, cost-aware pipeline design

Preferred skills

AWS Trainium/Inferentia (Neuron SDK), NeuronCore architecture, compiler behavior understanding

Technologies

Amazon SageMaker, AWS Trainium, PyTorch, Neuron SDK

Responsibilities

Write and optimize PyTorch training code for Trainium hardware, diagnose hardware-specific training issues, tune distributed training for throughput and cost, translate training requests into end-to-end production pipelines, collaborate with client teams on production workloads

Seniority

Staff, hands-on IC with client-facing delivery

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.