Staff Debug and DevOps Engineer
Core
Building and operating infrastructure for engineering teams to ensure reliable, scalable environments for AI hardware and software development.
Role type
Staff Debug and DevOps Engineer
Builds
CI/CD pipelines, automation frameworks, monitoring, logging, and observability systems for AI hardware/software stacks.
Domain
AI hardware, custom silicon (RISC-V), distributed systems, infrastructure
Deliverable
infrastructure
Required skills
Linux systems administration, troubleshooting distributed applications, CI/CD pipeline management, automation scripting (Python, Bash, Go), observability platform expertise (Prometheus, Grafana, OpenTelemetry, ELK), root-cause analysis
Preferred skills
Experience in Site Reliability Engineering, production engineering, complex distributed environments
Technologies
RISC-V, Python, Bash, Go, Prometheus, Grafana, OpenTelemetry, ELK, CI/CD tools
Responsibilities
Investigate and resolve complex issues across hardware and software systems; Build and maintain CI/CD infrastructure and automation frameworks; Design and enhance monitoring, logging, and alerting systems; Partner with software, firmware, and silicon teams to diagnose multi-layer issues; Drive long-term reliability improvements through automation.
Seniority
Staff, hands-on IC with strategic impact