Software Development Engineer, AI/ML, AWS Neuron, Model Inference
Core
Architect and implement distributed inference support for PyTorch on AWS custom ML accelerators (Trainium/Inferentia), optimizing LLMs like Llama and DeepSeek for latency and throughput.
Role type
Senior IC machine-learning engineer (inference acceleration)
Builds
High-performance kernels, distributed inference solutions, and system-level optimizations for the AWS Neuron SDK
Domain
Cloud infrastructure, machine learning, high-performance computing, custom silicon
Deliverable
production ML models
Required skills
C++, Python, system-level programming, ML model optimization, distributed computing, performance profiling, hardware-software co-design, fusion/sharding/tiling/scheduling
Preferred skills
PyTorch, JIT compilation, A
Technologies
PyTorch, JAX, AWS Neuron SDK, Trainium, Inferentia
Responsibilities
Design and optimize ML models/frameworks for custom hardware; build infrastructure to onboard diverse model architectures; implement high-performance kernels; analyze system-level performance bottlenecks; conduct comprehensive testing and continuous deployment; collaborate with compiler/runtime teams; work directly with customers on model enablement.
Seniority
Senior, hands-on IC with mentorship responsibilities