CareerPlanSign in

Site Reliability Engineer, Inference Infrastructure

Toronto🌐 Remote💼 Full-time🗓 2026-01-12 → 2026-09-26

Core

Build high-performance, scalable, and reliable machine learning systems for deploying frontier AI models via API endpoints.

Role type

Site Reliability Engineer (Inference Infrastructure)

Builds

AI platform delivering large language models through API endpoints

Domain

Artificial Intelligence / Large Language Models / Cloud Infrastructure

Deliverable

production ML models

Required skills

Kubernetes, distributed systems, Linux, GCP/Azure/AWS/OCI, GPU/TPU management, Golang/C++, cost management, SLO definition

Preferred skills

Custom Kubernetes operators, multi-cloud/hybrid serving, complex Linux troubleshooting, accelerator computational characteristics

Technologies

Kubernetes, GCP, Azure, AWS, OCI, Golang, C++

Responsibilities

Build self-service systems for managing and deploying services; Automate environment observability and resilience; Participate in on-call rotation to ensure SLOs; Develop team through knowledge sharing and code reviews

Seniority

Mid-Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.