Software Engineer, Kernel Reliability
Core
Improving the reliability of advanced compute clusters and underlying inference/training services for AI workloads.
Role type
Senior IC software engineer (kernel reliability & systems)
Builds
Internal and customer-facing AI compute clusters and production services
Domain
AI hardware infrastructure / Systems programming
Deliverable
production ML models | infrastructure
Required skills
C/C++, Python, operating systems, computer architecture, systems programming, debugging, root-cause analysis
Preferred skills
parallel/distributed programming, debug tool development, distributed application debugging, computer architecture concepts, incident response
Technologies
C, C++, Python, debuggers, core dumps, sanitizers, profilers, tracing
Responsibilities
Contribute to technical roadmap for kernel-centric reliability; Partner with operations to reduce downtime via tooling and debugging; Enhance debug tools for faster failure analysis; Collaborate on software stack and kernel improvements; Co-design next-gen architectures with hardware teams; Participate in incident response and post-mortems
Seniority
Senior, hands-on IC