Director Site Reliability Engineer, AI Infrastructure
Core
Build and lead d-Matrix's Site Reliability Engineering function from scratch, owning infrastructure for AI inference silicon development, validation, and customer-facing deployments across colocation, on-premises, and cloud.
Role type
Director, Site Reliability Engineering
Builds
Reliable and scalable infrastructure for AI inference silicon engineering and customer platform services
Domain
AI Infrastructure / Hardware Silicon / Hybrid Cloud
Deliverable
infrastructure
Required skills
SRE function building, Linux systems operations, Infrastructure as Code, Kubernetes cluster operations, observability stack ownership, Python/Go scripting, executive communication, high-ambiguity environment navigation
Preferred skills
Customer-facing infrastructure operations, high-speed interconnect fabrics (InfiniBand/RoCE/NVLink), HPC job schedulers (Slurm/LSF), multi-cloud hybrid operations, FinOps expertise, ITIL frameworks
Technologies
Terraform, Ansible, Prometheus, Grafana, Datadog, Splunk, Kubernetes, AWS, Azure, GCP, InfiniBand, RoCE, NVLink, Slurm, LSF
Responsibilities
Define SRE charter and hire team, direct Data Center & Lab Technician team, own 24x7 reliability across environments, establish SRE processes (SLOs, on-call, incident management), own observability stack, drive IaC automation and self-healing, manage FinOps and capacity planning, lead shared storage platform migration
Seniority
Director, hands-on IC with team leadership