CareerPlanGet AI match score →

Senior Software Engineer- ML Network Stack, ML Network Stack - Annapurna Labs

Tel Aviv-Yafo, Tel Aviv, Israel💼 Full-time🗓 2026-06-14 → 2026-07-24

Core

Building and maintaining infrastructure to monitor and report on functionality and performance of massive testing workloads for ML and HPC systems.

Role type

Senior Software Engineer (Infrastructure & Observability)

Builds

Automated monitoring, alerting, and performance reporting systems for large-scale ML/HPC clusters

Domain

Cloud Infrastructure, High-Performance Computing (HPC), Machine Learning Systems

Deliverable

production ML models | infrastructure

Required skills

Linux system administration, Python programming, CI/CD automation, performance data analysis, software architecture design, SW/HW co-design

Preferred skills

Experience with Grafana dashboards, NCCL/NVSHMEM frameworks, embedded systems, high-speed networking (RDMA)

Technologies

Python, Linux, AWS Managed Grafana, AWS Athena, NCCL, NVSHMEM, NIXL

Responsibilities

Automate software delivery to customers using internal CI/CD tools, write Python code to spool up large clusters and run benchmarks, digest performance data to create dashboards, invent automatic alerting mechanisms for regressions, manage complexity across many instance types and software stacks

Seniority

Senior, hands-on IC with mentorship responsibilities

Rewrite
## Responsibilities - Be a senior engineer on a team that builds and maintains the infrastructure that monitors and reports on functionality and performance of massive testing workloads run at scale. - Use internal Amazon CI/CD tools, Linux, and public AWS products to automate the delivery of our software to customers, saving developer time. - Write Python code that effortlessly spools up large clusters and runs benchmarks and applications for ML and HPC workloads. - Use AWS Managed Grafana and Athena to digest the massive amount of performance data generated by these workloads and create dashboards for developers and stakeholders. - Invent automatic mechanisms to alert developers to functional and performance regressions so they never reach customers. - Manage the complexity of infrastructure that covers many instance types, software stacks, Linux operating systems, cutting-edge releases and make it easy to evolve. ## Requirements - 5+ years of non-internship professional software development experience - 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience - 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience - 3+ years as a mentor, tech lead or leading engineering teams - 3+ years experience in SW/HW Co-Design - Solid knowledge of Linux, networking, and performant coding - Experience with embedded systems - Experience with high-speed networking or HPC/RDMA interconnects ## Nice to Have - Bachelor's degree in computer science or equivalent - Experience creating automated dashboards and visualization (such as Grafana) ## Benefits - Work/Life Balance - Mentorship & Career Growth - Inclusive culture - Flexible working culture - Opportunities for career advancement - Support for diverse experiences - Access to career-advancing resources ## About the Team The organization you would be joining is Annapurna Labs, an integral part of AWS that develops hardware and software components that are critical building blocks for EC2 infrastructure. Every instance in EC2 is running some type of hardware designed by Annapurna Labs. We specialize in designing software, systems, and chips that optimize the AWS customer experience.
Sourced via amazon · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply at Amazon ↗