Senior System Software Engineer - GPU Performance
Core
Performance engineer optimizing communication libraries (NCCL, NVSHMEM, UCX) for Deep Learning and HPC applications across large-scale GPU clusters.
Role type
Senior System Software Engineer (GPU Performance)
Builds
High-performance communication libraries for AI and HPC workloads
Domain
High Performance Computing (HPC) / Artificial Intelligence / GPU Systems
Deliverable
production ML models | infrastructure
Required skills
Parallel programming, Communication runtimes (MPI, NCCL, UCX, NVSHMEM), Performance benchmarking, Systems software fundamentals, C/C++ implementation, Debugging HW/SW stack, Python scripting, Container orchestration (Kubernetes, SLURM, Ansible, Docker)
Preferred skills
Infiniband/Ethernet networking (RDMA, congestion control), CUDA programming, Deep Learning Frameworks (PyTorch, TensorFlow)
Responsibilities
Conduct performance characterization on multi-GPU/multi-node clusters, Analyze HW-SW stack interactions, Evaluate proof-of-concepts and trade-offs, Triage and root-cause performance issues, Build tools for performance data visualization
Seniority
Senior, hands-on IC
