CareerPlanGet AI match score →

Site Reliability Engineer, Inference Infrastructure

Toronto, Ontario, Canada🌐 Remote💼 Full-time🗓 2026-06-11 → 2026-07-27

Required skills

5+ years of engineering experience running production infrastructure at a large scale, Experience designing large, highly available distributed systems with Kubernetes, and GPU workloads on those clusters, Experience with Kubernetes dev and production coding and support, Experience with GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid serving, Experience in designing, deploying, supporting, and troubleshooting in complex Linux-based computing environments, Experience in compute/storage/network resource and cost management, Strong understanding or working experience with distributed systems, Experience in Golang, C++ or other languages designed for high-performance scalable servers.

Preferred skills

Familiarity with computational characteristics of accelerators (GPUs, TPUs, and/or custom accelerators), especially how they influence latency and throughput of inference.

Technologies

Kubernetes, GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid serving, Linux-based computing environments, Golang, C++.

Responsibilities

Build self-service systems that automate managing, deploying and operating services. This includes our custom Kubernetes operators that support language model deployments, Automate environment observability and resilience, Enable all developers to troubleshoot and resolve problems, Take steps required to ensure we hit defined SLOs, including participation in an on-call rotation, Build strong relationships with internal developers and influence the Infrastructure team’s roadmap based on their feedback, Develop our team through knowledge sharing and an active review process.

Seniority

Not specified.

Domain

AI, Machine Learning, NLP, Infrastructure, Cloud Computing.

Sourced via linkedin · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on LinkedIn ↗