Sr Software Dev Engineer, SageMaker Training
Core
Build distributed services for large-scale training and reinforcement learning workloads, enabling customers to customize foundation models at scale.
Role type
Senior Software Development Engineer (Distributed Systems/ML Infrastructure)
Builds
Scheduling, capacity, and fleet-health systems for GPU clusters; multi-tenant GPU sharing infrastructure; engines for post-training loops.
Domain
Cloud Infrastructure / Machine Learning / Distributed Systems
Deliverable
production ML models
Required skills
Distributed systems design, GPU cluster management, reinforcement learning workload orchestration, multi-tenant architecture, system reliability, architectural leadership, mentoring
Preferred skills
Full SDLC ownership, code review leadership, source control management, build processes, testing strategies, operations experience
Technologies
GPU clusters, reinforcement learning engines, distributed scheduling systems
Responsibilities
Design and operate distributed services for large-scale training and RL workloads; build scheduling and fleet-health systems for GPU clusters; develop multi-tenant GPU sharing infrastructure; partner with Science and Product teams to translate requirements into scalable solutions; maintain operational excellence for critical training workloads.
Seniority
Senior, hands-on IC with architectural leadership