CareerPlanSign in

Software Development Engineer AI/ML, Inference Model Enablement, AWS Neuron

Cupertino, California, United States💼 Full-time💰 $193,300–$261,500🗓 2026-07-29 → 2026-09-26

Core

Onboard and optimize state-of-the-art open-source and customer LLMs (dense and MoE) for inference on Trainium accelerators, while advancing inference usability and quality through features, infrastructure optimization, tools, and automation.

Role type

Senior IC software development engineer (inference model enablement)

Builds

High-performance inference models and the Neuron serving stack for AWS Trainium

Domain

Cloud-scale machine learning infrastructure and large language model inference

Deliverable

production ML models

Required skills

distributed inference libraries, performance optimization, system reliability, architectural design, mentoring, technical leadership, full software development lifecycle, object-oriented design

Preferred skills

Machine Learning fundamentals, Large Language Model architecture, training/inference lifecycles, model execution optimization

Technologies

AWS Neuron, Trainium, vLLM, C++, Java, C#

Responsibilities

Deliver high-performance models using distributed inference libraries, drive technical excellence in performance optimization and system reliability, mentor team members, drive architectural decisions, collaborate with cross-functional teams to define technical strategy, author technical documentation and design proposals

Seniority

Senior, hands-on IC with technical leadership

Sourced via amazon · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.