CareerPlanSign in

Site Reliability Engineer

Toronto🌐 Remote💼 Full-time🗓 2026-07-14 → 2026-09-26

Core

Design, operate, and improve reliable infrastructure for large-scale AI training and inference workloads, including high-performance networks, GPU clusters, and storage.

Role type

Senior Site Reliability Engineer (AI Infrastructure)

Builds

Production-grade AI systems for natural, capable, and useful human-AI communication

Domain

Artificial Intelligence / High-Performance Computing / Distributed Systems

Deliverable

production ML models | infrastructure

Required skills

Linux administration, scripting, networking (BGP, InfiniBand), cluster scheduling (Kubernetes, SLURM), distributed storage (Ceph), GPU administration, incident response, automation, capacity planning

Preferred skills

NVIDIA GPU/CUDA/NCCL expertise, high-performance interconnects (RDMA, RoCE), Terraform/Ansible, observability (Prometheus, Grafana), bare-metal automation, cloud infrastructure (AWS, GCP, Azure)

Technologies

Kubernetes, SLURM, MAAS, Ceph, InfiniBand, CUDA, Prometheus, Grafana, Terraform, Ansible, AWS, GCP, Azure

Responsibilities

Design and operate infrastructure for AI training/inference; automate operational workflows; build monitoring and incident-response practices; diagnose performance and reliability issues; partner with ML/research teams; improve provisioning and deployment automation; plan cluster growth and lifecycle management

Seniority

Senior, hands-on IC

Sourced via lever · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.