CareerPlanSign in

Distributed Software Engineer

Toronto, CAN💼 Full-time🗓 2026-09-08 → 2026-09-26

Core

Building and operating production distributed systems and infrastructure software for clusters of Cerebras AI chips, servers, and switches.

Role type

Senior IC distributed systems engineer (infrastructure)

Builds

Cloud-scale clusters of wafer-scale AI systems, Kubernetes operators, and control-plane services

Domain

AI hardware infrastructure, distributed systems, cloud operations

Deliverable

production ML models | infrastructure

Required skills

Go, Python, Kubernetes (operators, CRDs, reconciliation), distributed systems debugging, Linux, networking, Prometheus, Grafana

Preferred skills

bare-metal/HPC fleet operations, scheduler internals, RDMA/RoCE, eBPF, Ceph, NVMe-oF, etcd

Technologies

Go, Python, Kubernetes, gRPC, Prometheus, Grafana, Redfish, IPMI, gNMI, sFlow

Responsibilities

Automate bare-metal networking, OS, and application software across clusters; implement Kubernetes operators for large inference workloads; build gRPC control-plane services with authorization and quota policy; design metrics and log pipelines with custom exporters; implement failure detection, HA control planes, and automated recovery; develop CLIs, APIs, and MCP gateways for fleet access.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.