Principal Debug and SRE Lead
Core
Leading a team to ensure reliability, observability, and operational excellence of infrastructure supporting AI hardware and software development.
Role type
Principal Debug & Site Reliability Engineering Lead
Builds
Debugging methodologies, monitoring systems, automation, and engineering workflows for AI compute clusters
Domain
AI hardware, custom silicon (RISC-V), distributed infrastructure, firmware, and operating systems
Deliverable
infrastructure
Required skills
Linux systems debugging, root-cause analysis, automation, technical leadership, cross-functional collaboration
Preferred skills
Mentoring engineers, driving technical execution, improving operational excellence
Technologies
Python, C++, Go, Bash, Prometheus, Grafana, OpenTelemetry, ELK
Responsibilities
Lead team responsible for reliability and operational health of engineering infrastructure; Drive root-cause analysis of complex issues spanning silicon, firmware, OS, and networking; Build and improve debugging methodologies and automation; Partner with silicon, firmware, and software teams to resolve critical issues; Mentor engineers and establish technical direction
Seniority
Principal, hands-on IC with leadership