Senior Software Engineer- ML Network Stack, ML Network Stack - Annapurna Labs
Core
Building and maintaining infrastructure to monitor and report on functionality and performance of massive testing workloads for ML and HPC systems.
Role type
Senior Software Engineer (Infrastructure & Observability)
Builds
Automated monitoring, alerting, and performance reporting systems for large-scale ML/HPC clusters
Domain
Cloud Infrastructure, High-Performance Computing (HPC), Machine Learning Systems
Deliverable
production ML models | infrastructure
Required skills
Linux system administration, Python programming, CI/CD automation, performance data analysis, software architecture design, SW/HW co-design
Preferred skills
Experience with Grafana dashboards, NCCL/NVSHMEM frameworks, embedded systems, high-speed networking (RDMA)
Technologies
Python, Linux, AWS Managed Grafana, AWS Athena, NCCL, NVSHMEM, NIXL
Responsibilities
Automate software delivery to customers using internal CI/CD tools, write Python code to spool up large clusters and run benchmarks, digest performance data to create dashboards, invent automatic alerting mechanisms for regressions, manage complexity across many instance types and software stacks
Seniority
Senior, hands-on IC with mentorship responsibilities