CareerPlanGet AI match score →

Director, Site Reliability Engineering - AI Accelerator Infrastructure - Contract

Santa Clara💼 Contract🗓 2026-07-13 → 2026-07-31

Core

Building and leading a Site Reliability Engineering function from scratch to own the physical and virtual infrastructure layer for a purpose-built AI inference silicon company.

Role type

Director, Site Reliability Engineering (Infrastructure)

Builds

Reliable, scalable infrastructure for development, validation, and customer-facing deployments of AI inference silicon across colocation, on-premises, and cloud.

Domain

AI Hardware / Semiconductor / Infrastructure

Deliverable

infrastructure

Required skills

SRE discipline establishment, team leadership (3-5 engineers), Linux systems expertise, hybrid-cloud storage integration, colocation operations, Infrastructure as Code, Kubernetes operations, FinOps, incident management, observability stack ownership.

Preferred skills

AIOps, enterprise shared storage platform migration, follow-the-sun on-call model design.

Technologies

Terraform, Ansible, Prometheus, Grafana, Datadog, AWS, Azure, GCP, NAS/SAN, NFS/SMB.

Responsibilities

Define SRE charter and technical roadmap; hire and grow the SRE team; establish SLOs, error budgets, and incident management frameworks; own 24x7 reliability and observability; drive FinOps and capacity planning; lead migration to enterprise-grade shared storage; direct data center and lab technician teams.

Seniority

Director, hands-on engineering leader

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗