Senior Systems Software Engineer - GPU Performance at Scale
Core
Lead performance engineering for large-scale GPU infrastructure, aligning AI workloads with datacenter builds, and optimizing high-performance compute platforms.
Role type
Senior Systems Software Engineer (GPU Performance)
Builds
Large-scale performance platforms, tools, and methodologies for AI workloads on NVIDIA GPUs, CPUs, and networking hardware.
Domain
AI, High-Performance Computing (HPC), GPU Computing, Datacenter Infrastructure
Deliverable
production ML models | infrastructure
Required skills
CUDA, C/C++/Python/Bash, Linux systems programming, container technology (Docker), HPC/deep learning experience, performance analysis and debugging, systems architecture
Preferred skills
Slurm, virtualization, cloud platform solutions, scheduling and resource management, end-to-end GPU profiling
Technologies
CUDA, Slurm, Docker, Linux, C, C++, Python, Bash
Responsibilities
Implement performance practices in large-scale GPU infrastructure; Align next-generation AI workloads with datacenter builds; Develop engineering solutions for continuous performance insights; Decompose high-complexity performance or stability issues; Collaborate with SW/FW teams to develop methods and tools for resolving critical issues.
Seniority
Senior, hands-on IC