Lead Software Systems Engineer - GPU Performance
Core
Optimizing performance of large-scale GPU clusters at the intersection of hardware and software for a full-stack AI cloud platform.
Role type
Lead Software Systems Engineer (GPU Performance)
Builds
Hyperscaler platform components including GPU orchestration, inference optimization, and distributed communication layers.
Domain
Cloud Infrastructure / High-Performance Computing (HPC) / AI Hardware
Deliverable
production ML models
Required skills
System-level software development, Linux systems administration and tuning, Server architecture (PCIe, NICs, Kernel), Performance-oriented programming (C/C++, Go, Python)
Preferred skills
Hardware vendor collaboration, Cluster qualification, Troubleshooting real workloads
Technologies
InfiniBand, RoCE, KVM, QEMU, MPI, NCCL, NVIDIA, Mellanox, Intel
Responsibilities
Identify performance bottlenecks and drive improvements for cluster operation and tuning; Investigate and troubleshoot GPU cluster performance under training and inference workloads; Evaluate and integrate new hardware, configurations, and tuning approaches; Support complex performance-related escalations; Contribute to hardware and cluster qualification to ensure performance expectations are met
Seniority
Senior, hands-on IC