Director, Site Reliability Engineering - AI Accelerator Infrastructure - Contract
Core
Building and leading a Site Reliability Engineering function from scratch to own the physical and virtual infrastructure layer for a purpose-built AI inference silicon company.
Role type
Director, Site Reliability Engineering (Infrastructure)
Builds
Reliable, scalable infrastructure for development, validation, and customer-facing deployments of AI inference silicon across colocation, on-premises, and cloud.
Domain
AI Hardware / Semiconductor / Infrastructure
Deliverable
infrastructure
Required skills
SRE discipline establishment, team leadership (3-5 engineers), Linux systems expertise, hybrid-cloud storage integration, colocation operations, Infrastructure as Code, Kubernetes operations, FinOps, incident management, observability stack ownership.
Preferred skills
AIOps, enterprise shared storage platform migration, follow-the-sun on-call model design.
Technologies
Terraform, Ansible, Prometheus, Grafana, Datadog, AWS, Azure, GCP, NAS/SAN, NFS/SMB.
Responsibilities
Define SRE charter and technical roadmap; hire and grow the SRE team; establish SLOs, error budgets, and incident management frameworks; own 24x7 reliability and observability; drive FinOps and capacity planning; lead migration to enterprise-grade shared storage; direct data center and lab technician teams.
Seniority
Director, hands-on engineering leader