SDE, Neuron Inference, Neuron Inference
Core
Senior software engineer developing and optimizing core building blocks (Attention, MLP, Quantization, Speculative Decoding, MoE) for LLM inference on AWS Neuron chips.
Role type
Senior IC machine-learning inference software engineer
Builds
High-performance LLM inference applications on AWS Inferentia/Trainium accelerators
Domain
Cloud-scale machine learning inference, hardware acceleration
Deliverable
production ML models
Required skills
C++, C#, Java, Perl, Object Oriented Design, multi-threaded/distributed systems, system design, large-scale software development
Preferred skills
Full SDLC experience, code reviews, source control, build processes, testing, operations
Technologies
AWS Neuron, Inferentia, Trainium, LLM frameworks
Responsibilities
Adapt latest LLM optimization research to Neuron chips, extract best performance from open source and internal models, collaborate with chip architects and compiler engineers
Seniority
Senior, hands-on IC