CareerPlanSign in
This role has closed. Browse similar live roles below, or see all open jobs →

Lead Site Reliability Engineer

TORONTO, Ontario, Canada💼 Full-time🗓 2026-09-10 → 2026-09-25

Core

Lead the production environment for distributed software applications, ensuring health, security, and availability through monitoring, tooling, and operational support.

Role type

Senior Site Reliability Engineer (SRE)

Builds

Production-grade monitoring tools, alerting systems, and automated infrastructure for large distributed applications.

Domain

Financial Services / Cloud Native / Distributed Systems

Deliverable

production ML models | product features | dashboards & analysis | infrastructure

Required skills

Kubernetes, OpenShift, Azure, Java/SpringBoot, Python, SQL/No-SQL, Microservices, Container Orchestration, System Monitoring, Performance Tuning, Capacity Planning

Preferred skills

Mainframes & JCL, Multi-cloud/Hybrid environments, Docker, Splunk, Dynatrace, RUM, Grafana, OCP certification, GitHub

Technologies

Azure, OpenShift, Kubernetes, Docker, Splunk, Dynatrace, Grafana, RUM, Java, SpringBoot, Python, SQL, No-SQL

Responsibilities

Monitor system availability and holistic health; build tools for platform infrastructure management; debug production issues across the stack; drive tool creation for health monitoring and alerting; optimize system performance and reliability; analyze metrics for performance tuning and fault finding; participate in system design consulting and capacity planning; create sustainable systems through automation.

Seniority

Senior, hands-on IC with mentorship responsibilities

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.