CareerPlanSign in

Director Site Reliability Engineer, AI Infrastructure

Santa Clara💼 Full-time🗓 2026-09-26 → 2026-09-27

Core

Build and lead d-Matrix's Site Reliability Engineering function from scratch, owning infrastructure for AI inference silicon development, validation, and customer-facing deployments across colocation, on-premises, and cloud.

Role type

Director, Site Reliability Engineering

Builds

Reliable and scalable infrastructure for AI inference silicon engineering and customer platform services

Domain

AI Infrastructure / Hardware Silicon / Hybrid Cloud

Deliverable

infrastructure

Required skills

SRE function building, Linux systems operations, Infrastructure as Code, Kubernetes cluster operations, observability stack ownership, Python/Go scripting, executive communication, high-ambiguity environment navigation

Preferred skills

Customer-facing infrastructure operations, high-speed interconnect fabrics (InfiniBand/RoCE/NVLink), HPC job schedulers (Slurm/LSF), multi-cloud hybrid operations, FinOps expertise, ITIL frameworks

Technologies

Terraform, Ansible, Prometheus, Grafana, Datadog, Splunk, Kubernetes, AWS, Azure, GCP, InfiniBand, RoCE, NVLink, Slurm, LSF

Responsibilities

Define SRE charter and hire team, direct Data Center & Lab Technician team, own 24x7 reliability across environments, establish SRE processes (SLOs, on-call, incident management), own observability stack, drive IaC automation and self-healing, manage FinOps and capacity planning, lead shared storage platform migration

Seniority

Director, hands-on IC with team leadership

Sourced via ashby · Listed on CareerPlan, which tracks 833,000+ jobs from 20+ sources.