CareerPlanGet AI match score →

Ingénieur fiabilité des infrastructures

Montreal, Quebec, Canada🌐 Remote💼 Full-time🗓 2025-12-16 → 2026-07-31

Core

Maintaining, optimizing, and ensuring reliability and performance of critical SaaS infrastructure on AWS and Kubernetes with a focus on automation, observability, and continuous improvement.

Role type

Senior Site Reliability Engineer (SRE)

Builds

Production SaaS platforms on AWS and Kubernetes for healthcare supply chain clients

Domain

Cloud Infrastructure / Healthcare Supply Chain

Deliverable

production ML models | product features | dashboards & analysis | infrastructure

Required skills

Incident management, post-incident analysis (RCA), SLO/SLI definition, observability tooling, infrastructure as code (IaC), CI/CD pipelines, system design consultation, technical documentation, cross-team coordination

Preferred skills

None stated

Technologies

AWS, Kubernetes, Datadog, Terraform, GitLab CI/CD

Responsibilities

Collaborate with engineering teams on pre-launch system design and capacity planning; Identify weaknesses and lead initiatives to simplify and strengthen the platform; Monitor availability, latency, and system health; Optimize observability by defining SLO/SLI and creating actionable dashboards; Develop automation tools, IaC frameworks, and CI/CD pipelines to reduce manual intervention; Implement sustainable incident management and lead post-incident reviews (RCA); Act as incident commander during incidents to coordinate response and ensure rapid restoration

Seniority

Senior, hands-on IC

Sourced via workable · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workable ↗