CareerPlanSign in

Software Engineer II - AI/ML, Neuron Inference

Cupertino, California, United States💼 Full-time🗓 2026-09-23 → 2026-09-25

Core

Design, develop, and optimize machine learning models and frameworks for deployment on custom ML hardware accelerators (AWS Trainium), focusing on distributed inference and high-performance kernels.

Role type

Senior IC software engineer (AI/ML inference optimization)

Builds

Distributed inference solutions for PyTorch and large-scale LLMs on AWS Trainium

Domain

Cloud infrastructure, Machine Learning, High-Performance Computing, Custom Hardware Acceleration

Deliverable

production ML models

Required skills

Python, C++, distributed computing, system-level programming, ML model optimization, performance profiling, hardware architecture knowledge, parallel computing, memory management

Preferred skills

CUDA kernel development, Triton syntax, vLLM, SGLang, TensorRT, JIT compilation, computer architecture

Technologies

PyTorch, AWS Trainium, Neuron SDK, CUTLASS, FlashInfer

Responsibilities

Design and implement high-performance kernels for ML operations; optimize system-level performance across Neuron hardware generations; build infrastructure to onboard diverse model architectures; conduct performance analysis and resolve bottlenecks; implement optimizations like fusion, sharding, and tiling; collaborate with compiler and runtime teams; work directly with customers to enable model performance.

Seniority

Senior, hands-on IC

Sourced via amazon · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.