Senior Systems Software Engineer, GPU Compute
Core
Developing and optimizing core components of a hyperscaler AI cloud platform, specifically focusing on GPU computing, InfiniBand networks, and the KVM/QEMU stack to support high-performance HPC environments.
Role type
Senior Systems Software Engineer (GPU Compute & Virtualization)
Builds
High-performance GPU clusters, InfiniBand fabrics, and virtualized HPC infrastructure for AI/ML workloads.
Domain
Cloud Infrastructure, High-Performance Computing (HPC), GPU Computing, Virtualization
Deliverable
production ML models | infrastructure
Required skills
System-level software development, Linux systems administration, Server architecture (PCIe, NICs, Kernel), Performance-oriented programming (C/C++, Go, Python)
Preferred skills
GPU end-to-end testing in cluster environments, HPC workload optimization, RDMA/RoCE/InfiniBand protocols, Software-Defined Networking (SDN), QEMU/KVM virtualization, Deep learning frameworks (PyTorch, TensorFlow), Collective communication libraries (MPI, NCCL)
Technologies
Kubernetes, QEMU, KVM, InfiniBand, MPI, NCCL, PyTorch, TensorFlow
Responsibilities
Tuning performance of GPU clusters and InfiniBand networks, Analyzing and troubleshooting root causes of GPU/InfiniBand issues, Integrating new GPU hardware into infrastructure, Enhancing automation systems for proactive monitoring and fault resolution, Configuring and managing GPU devices and InfiniBand fabrics
Seniority
Senior, hands-on IC