CareerPlanSign in

Software Engineer-AI/ML, Inference Team - AWS Neuron

Seattle, Washington, United States💼 Full-time💰 $143,700–$194,400🗓 2026-10-06 → 2026-10-07

Core

Build serving technology for large-scale model inference on AWS Neuron chips, optimizing performance within open-source frameworks like vLLM and SGLang.

Role type

Senior IC software engineer (ML inference)

Builds

Open-source inference frameworks (vLLM, SGLang) and production-ready model serving tooling

Domain

Cloud-scale machine learning inference, AI accelerators

Deliverable

production ML models

Required skills

C++, C#, Java, Perl, distributed systems, multi-threaded software, transformer architecture, LLM fundamentals, model optimization

Preferred skills

Kernel development (CUDA, Triton), LLM deployment on AI hardware, profiling large-scale systems, vLLM/SGLang/TensorRT production experience

Technologies

AWS Neuron, vLLM, SGLang, CUDA, Triton, TensorRT (via careerplan.io/jobs/10569950-software-engineer-aiml-inference-team-aws-neuron-at-amazon)

Responsibilities

Implement and upstream support for continuous batching, paged attention, quantization, and distributed inference; build automation for model deployment; improve test and benchmarking workflows; collaborate on end-to-end model performance; apply engineering practices for reliable inference.

Seniority

Senior, hands-on IC