Software Development Engineer I – AI/ML Network Infrastructure, Annapurna Labs
Core
Develop high-performance C/C++ code for network communication libraries (NCCL, NVSHMEM, NIXL) enabling distributed AI/ML training across massive GPU clusters.
Role type
Software Development Engineer I – AI/ML Network Infrastructure
Builds
Network stack and communication frameworks for EC2 distributed AI/ML systems
Domain
Cloud infrastructure, High-Performance Computing (HPC), Machine Learning
Deliverable
production ML models
Required skills
C/C++, Linux internals, Parallel Computer Architecture, Distributed Systems, Performance profiling
Preferred skills
RDMA/high-speed interconnects, GPU programming (CUDA), Hardware-software co-design, Open-source contributions
Technologies
NCCL, NVSHMEM, NIXL, Python, AWS tools (CI/CD, Grafana, Athena), CUDA
Responsibilities
Write high-performance C/C++ code for network communication libraries, Build and maintain infrastructure monitoring large-scale AI/ML workloads, Develop automation for testing and benchmarking, Design mechanisms to detect functional and performance regressions, Work across many instance types and Linux environments
Seniority
Early-career IC