CareerPlanGet AI match score →

Engineering Manager, Kernel Reliability

US and Canada Offices💼 Full-time🗓 2026-01-08 → 2026-08-01

Core

Leading a team to improve the reliability of advanced compute clusters and underlying inference/training services for AI workloads.

Role type

Senior Engineering Manager, Kernel Reliability

Builds

Internal and customer-facing AI compute clusters and software services

Domain

AI hardware (ASIC) and distributed systems reliability

Deliverable

production ML models | infrastructure

Required skills

parallel and distributed programming, debug and diagnostic tool development, failure analysis, computer architecture, monitoring and reliability engineering, team leadership

Preferred skills

experience with GPU/embedded systems, incident response, post-mortem analysis

Technologies

debuggers, core dump handling, code sanitizers, multicore, message passing

Responsibilities

Own technical vision and roadmap for kernel-centric reliability; provide tooling and manual intervention for failure analysis; enhance debug tools; collaborate on software stack improvements; codesign next-gen architectures with reliability in mind; lead and mentor a high-caliber team

Seniority

Senior, hands-on IC with team leadership

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗