889 open positions
Sr Software Dev Engineer, SageMaker Training
AmazonBellevue, Washington, United States
$168k–$227k2d
Build distributed services for large-scale training and reinforcement learning workloads, enabling customers to customize foundation models at scale.
Engineering Manager, GPU Reliability, Accelerators↗
GoogleSunnyvale, CA, US
$207k–$207k4d
Lead the GPU Reliability Systems Operations and Tooling software organization to ensure the reliability of Google's massive GPU supercomputer systems used for AI advancements.
Senior Performance Engineer, Efficiency Red Team
AmazonSeattle, Washington, United States
4d
Hunting for hidden waste across the world's largest compute infrastructure to turn findings into freed capacity and cost savings.
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training
AmazonCupertino, California, United States
4d
Architect and implement business-critical features for distributed training on AWS Trainium, optimizing throughput and convergence for frontier-scale models.
Software Engineer- AI/ML, Amazon Neuron Training
AmazonCupertino, California, United States
$165k–$224k4d
Building distributed training infrastructure, parallelism techniques, and high-performance kernels for large-scale pretraining, post-training, and reinforcement learning workloads on AWS Trainium custom ML accelerators.
Sr. Software Engineer- AI/ML, Amazon Neuron Training
AmazonCupertino, California, United States
$193k–$262k4d
Lead the effort to build distributed training and post-training support for PyTorch and JAX on AWS Trainium accelerators, enabling large-scale training, post-training, and reinforcement learning workloads.
← Select a job to preview