Sr. Software Development Engineer, Inference Team - AWS Neuron
Core
Lead development of core serving technologies within open-source inference frameworks (vLLM, SGLang) to enable efficient large-scale model inference on AWS Neuron accelerators.
Role type
Senior IC software development engineer (LLM inference systems)
Builds
High-performance model inference solutions for customer workloads on AWS Neuron-powered instances
Domain
Cloud-scale machine learning / AI accelerator hardware
Deliverable
production ML models
Required skills
LLM serving performance optimization, kernel development, parallel computation, distributed KV cache, speculative decoding, systems engineering, open-source framework customization, design leadership, mentorship
Preferred skills
PyTorch or JAX development, LLM deployment on AI accelerators, vLLM or SGLang contribution, CUDA or Triton kernel development
Technologies
AWS Neuron, vLLM, SGLang, PyTorch, JAX, CUDA, Triton
Responsibilities
Customize and optimize open-source inference frameworks for AWS Neuron; lead design and architecture of new and existing systems; influence technical roadmap by evaluating emerging inference research; collaborate with model, compiler, runtime, and performance engineering teams
Seniority
Senior, hands-on IC with design leadership