Software Development Engineer, Alexa Excellence, Alexa LLM Inference, Capacity, & Efficiency
Core
Design, develop, and maintain large-scale distributed systems and infrastructure services for LLM inference, focusing on GPU fleet management, optimization, and cost efficiency for Alexa.
Role type
Senior IC software development engineer (LLM inference infrastructure)
Builds
High-visibility, high-impact systems for Alexa's next generation of AI experiences
Domain
Cloud infrastructure + Generative AI
Deliverable
production ML models
Required skills
distributed systems design, GPU fleet management, low-latency ML inference optimization, system architecture, code review, cross-functional collaboration
Preferred skills
full SDLC experience, LLM deployment on GPUs/Neuron/TPU, operational excellence for Tier-1 services
Technologies
GenAI, GPUs, Neuron, TPU, AWS Bedrock
Responsibilities
Design and maintain large-scale distributed systems; optimize GPU workloads for low-latency inference; collaborate on feature delivery from design to production; set standards for operational excellence; write clean, testable code; research performance and cost efficiency improvements
Seniority
Senior, hands-on IC