Sr. Manager, Software Development, ML Network Stack - Annapurna Labs
Core
Managing the software development team responsible for the network stack of EC2 distributed AI/ML and HPC systems, including the CI/CD framework for EFA ML and HPC Host stack components.
Role type
Senior Engineering Manager (Systems/Network)
Builds
Network stack functions and features for large-scale ML workloads on EC2
Domain
Cloud Infrastructure, High-Performance Computing (HPC), Distributed AI/ML
Deliverable
production ML models | product features
Required skills
Systems programming, Network protocols, RDMA, HPC, HW/SW co-design, Engineering team management, Agile project management, Consumer software development lifecycle
Preferred skills
NVIDIA stack experience, ML applications and frameworks, Large scale high-traffic application design
Technologies
RDMA, EFA, NVIDIA stack
Responsibilities
Direct work to ensure team delivery of functions and features for latest ML workloads, Manage multiple concurrent programs and development teams, Partner with product and program management teams
Seniority
Senior, hands-on IC with management responsibilities