CareerPlanSign in

Senior Software Engineer, GPU Infrastructure

CANADA🌐 Remote💼 Full-time🗓 2026-09-04 → 2026-09-26

Core

Build, deploy, and operate Kubernetes-based GPU/TPU superclusters to optimize AI/ML training performance, reliability, and cost.

Role type

Senior IC GPU infrastructure engineer

Builds

Scalable GPU/TPU clusters and self-service interfaces for distributed training workloads

Domain

Cloud infrastructure + AI/ML high-performance computing

Deliverable

infrastructure

Required skills

Kubernetes cluster management at scale, GPU/TPU cluster operations, distributed training frameworks (JAX, PyTorch), Python, Go, Linux internals, RDMA networking, performance optimization

Preferred skills

Open-source contribution experience

Technologies

Kubernetes, RDMA, NCCL, JAX, PyTorch

Responsibilities

Build and operate GPU/TPU superclusters across multiple clouds; optimize infrastructure for AI/ML training; diagnose and resolve infrastructure bottlenecks and failures; create self-service interfaces for researchers; collaborate with AI researchers and ML engineers; participate in 24x7 on-call rotation

Seniority

Senior, hands-on IC

Sourced via codingjobboard · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.