CareerPlanSign in

Sr. Specialist - Platform Operations (AI & Agentic Systems)

CA-Toronto-York St 24/25💼 Full-time💰 $91,000–$91,000🗓 2026-08-21 → 2026-09-26

Core

Manage, scale, and optimize production-grade Kubernetes clusters and AI/agentic systems to ensure reliable platform operations for global markets and clients.

Role type

Senior Infrastructure Engineer (Kubernetes & AI Operations)

Builds

Production-grade Kubernetes clusters, automated deployment pipelines, and operationalized AI/agentic services.

Domain

Financial Services / Cloud Infrastructure / AI Systems

Deliverable

production ML models | infrastructure

Required skills

Kubernetes administration, GitOps (ArgoCD), Infrastructure as Code (Terraform/OpenTofu/Pulumi), Observability (Prometheus/Grafana/ELK), Linux internals, Cloud architecture (AWS), Python/Go/Bash scripting

Preferred skills

Internal Developer Platforms (Backstage), Service meshes (Istio/Linkerd), MLOps, AI observability and cost optimization

Technologies

Kubernetes, ArgoCD, Terraform, OpenTofu, Pulumi, Prometheus, Grafana, ELK, OpenSearch, AWS, GitHub Actions, GitLab CI, Jenkins, Docker, containerd, Istio, Linkerd, Backstage

Responsibilities

Manage and optimize multi-cloud Kubernetes clusters; Design and maintain GitOps pipelines; Implement monitoring and alerting systems; Conduct post-mortems and automate operational toil; Collaborate on developer experience improvements; Deploy and monitor AI/agentic workflows; Implement automation for system testing and deployment.

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.