Software Engineering Manager - GPU Communications Libraries
Core
Lead the NVSHMEM and UCX communication libraries team to deliver high-performance networking solutions for Deep Learning and HPC applications running on tens of thousands of GPUs.
Role type
Senior Engineering Manager (HPC Networking & Systems Software)
Builds
Production-grade communication libraries (NCCL, NVSHMEM, UCX) and high-speed interconnect software for GPU clusters.
Domain
High-Performance Computing (HPC), Deep Learning Infrastructure, GPU Networking
Deliverable
production ML models | product features
Required skills
HPC networking or system software specialization, software management, systems software fundamentals, C/C++ programming, Linux debugging, project prioritization
Preferred skills
Parallel programming models (MPI, SHMEM), communication runtimes (NCCL, NVSHMEM, UCX), CUDA programming, RDMA technologies, Deep Learning frameworks (PyTorch, TensorFlow)
Technologies
NVLink, InfiniBand, Ethernet, CUDA, MPI, OpenMP, PyTorch, TensorFlow
Responsibilities
Lead and mentor the library engineering team, participate in feature design and implementation, collaborate with partners to define product roadmap, review and improve team processes and infrastructure
Seniority
Senior, hands-on IC with management responsibilities
