CareerPlanSign in

Staff, Reliability Engineer

Toronto💼 Full-time🗓 2026-07-31 → 2026-09-26

Core

Define reliability strategy for next-generation AI computing systems across data center and workstation products, ensuring high uptime and durability.

Role type

Staff Reliability Engineer (AI Hardware)

Builds

High-performance AI platforms (silicon, thermal, mechanical, electrical, software)

Domain

AI Hardware / High-Performance Computing

Deliverable

production ML models | product features | infrastructure

Required skills

Reliability engineering strategy, statistical analysis (HALT, HASS, ALT, MTBF, Weibull, FMEA), root-cause analysis, advanced cooling validation (vapor chambers, heat pipes, liquid cooling), cross-functional leadership, NPI process management, mentorship

Preferred skills

Experience in high-performance computing, AI hardware, or data center systems

Technologies

RISC-V, thermal lab equipment, liquid cooling systems

Responsibilities

Set reliability strategy for AI computing systems, lead root-cause investigations on failures, validate advanced cooling systems, act as point of contact for manufacturing partners, mentor engineers and lead design reviews

Seniority

Staff, strategic leadership with hands-on technical execution

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.