Engineering Manager, Kernel Reliability
Core
Leading a team to improve the reliability of advanced compute clusters and underlying inference/training services for AI workloads.
Role type
Senior Engineering Manager, Kernel Reliability
Builds
Internal and customer-facing AI compute clusters and software services
Domain
AI hardware (ASIC) and distributed systems reliability
Deliverable
production ML models | infrastructure
Required skills
parallel and distributed programming, debug and diagnostic tool development, failure analysis, computer architecture, monitoring and reliability engineering, team leadership
Preferred skills
experience with GPU/embedded systems, incident response, post-mortem analysis
Technologies
debuggers, core dump handling, code sanitizers, multicore, message passing
Responsibilities
Own technical vision and roadmap for kernel-centric reliability; provide tooling and manual intervention for failure analysis; enhance debug tools; collaborate on software stack improvements; codesign next-gen architectures with reliability in mind; lead and mentor a high-caliber team
Seniority
Senior, hands-on IC with team leadership