CareerPlanGet AI match score →

RPE - System Reliability Engineering Specialist (Hybrid)

Montreal, Canada💼 Full-time🗓 2026-07-02 → 2026-07-30

Core

Design, build, and maintain high-scale, reliable distributed systems and services for Morgan Stanley's technology platform, focusing on observability, automation, and troubleshooting across the full stack.

Role type

Senior System Reliability Engineering (SRE) Specialist

Builds

Scalable, reliable services and platforms for electronic trading, algorithm trading, cloud engineering, and infrastructure

Domain

Financial Services / Distributed Systems / Cloud Infrastructure

Deliverable

production ML models | product features | dashboards & analysis | infrastructure

Required skills

Distributed systems architecture, Linux/UNIX administration, Database management (DB2, Sybase, MongoDB), CI/CD pipelines, Scripting (Python, Bash, Perl), Observability tools (Grafana, Prometheus, Dynatrace, AppDynamics), Microservices, Load balancing, Queueing, Caching, Network stack debugging, On-call incident management

Preferred skills

Experience with data streaming technologies (Spark, Kafka), Three-tier architecture, Large-scale online systems operation

Technologies

Git, Artifactory, Jenkins, Docker, Spark, Kafka, DB2, Sybase, MongoDB, Grafana, Prometheus, Dynatrace, AppDynamics

Responsibilities

Design and maintain systems with engineering teams, troubleshoot hardware/software/application/network issues, drive automation for deployment and management, identify reliability risks, participate in design reviews and operational readiness, manage on-call rotation

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗