CareerPlanSign in

Software Engineer, GPU Infrastructure (HPC)

Canada🌐 Remote💼 Full-time🗓 2026-01-15 → 2026-09-25

Core

Building and operating GPU/TPU superclusters and HPC infrastructure to train, evaluate, and serve frontier AI models for developers and enterprises.

Role type

Staff Software Engineer (GPU Infrastructure)

Builds

Kubernetes-based GPU/TPU superclusters across multiple clouds for AI workloads

Domain

High-Performance Computing (HPC) / AI Infrastructure

Deliverable

infrastructure

Required skills

GPU/TPU cluster management, Kubernetes at scale, Python, Go, Linux internals, RDMA networking, distributed training frameworks (JAX, PyTorch, TensorFlow)

Preferred skills

Open-source contributions, observability, infrastructure-as-code (IaC), mentoring

Technologies

Kubernetes, RDMA, NCCL, JAX, PyTorch, TensorFlow

Responsibilities

Deploy and manage Kubernetes-based GPU/TPU superclusters; Optimize infrastructure for cost efficiency and performance using RDMA/NCCL; Troubleshoot complex infrastructure bottlenecks and failures; Design self-service tools for researchers; Collaborate with researchers on emerging ML needs; Advocate for observability and automation; Mentor team members through code reviews and documentation

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.