Software Engineer-AI/ML, Inference Team - AWS Neuron
Core
Build serving technology for large-scale model inference on AWS Neuron chips, optimizing performance within open-source frameworks like vLLM and SGLang.
Role type
Senior IC software engineer (ML inference)
Builds
Open-source inference frameworks (vLLM, SGLang) and production-ready model serving tooling
Domain
Cloud-scale machine learning inference, AI accelerators
Deliverable
production ML models
Required skills
C++, C#, Java, Perl, distributed systems, multi-threaded software, transformer architecture, LLM fundamentals, model optimization
Preferred skills
Kernel development (CUDA, Triton), LLM deployment on AI hardware, profiling large-scale systems, vLLM/SGLang/TensorRT production experience
Technologies
AWS Neuron, vLLM, SGLang, CUDA, Triton, TensorRT (via careerplan.io/jobs/10569950-software-engineer-aiml-inference-team-aws-neuron-at-amazon)
Responsibilities
Implement and upstream support for continuous batching, paged attention, quantization, and distributed inference; build automation for model deployment; improve test and benchmarking workflows; collaborate on end-to-end model performance; apply engineering practices for reliable inference.
Seniority
Senior, hands-on IC