CareerPlanSign in

Software Development Engineer, AI/ML, AWS Neuron, Model Inference

Cupertino, California, United States💼 Full-time🗓 2026-08-07 → 2026-09-26

Core

Architect and implement distributed inference support for PyTorch on AWS custom ML accelerators (Trainium/Inferentia), optimizing LLMs like Llama and DeepSeek for latency and throughput.

Role type

Senior IC machine-learning engineer (inference acceleration)

Builds

High-performance kernels, distributed inference solutions, and system-level optimizations for the AWS Neuron SDK

Domain

Cloud infrastructure, machine learning, high-performance computing, custom silicon

Deliverable

production ML models

Required skills

C++, Python, system-level programming, ML model optimization, distributed computing, performance profiling, hardware-software co-design, fusion/sharding/tiling/scheduling

Preferred skills

PyTorch, JIT compilation, A

Technologies

PyTorch, JAX, AWS Neuron SDK, Trainium, Inferentia

Responsibilities

Design and optimize ML models/frameworks for custom hardware; build infrastructure to onboard diverse model architectures; implement high-performance kernels; analyze system-level performance bottlenecks; conduct comprehensive testing and continuous deployment; collaborate with compiler/runtime teams; work directly with customers on model enablement.

Seniority

Senior, hands-on IC with mentorship responsibilities

Sourced via amazon · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.