CareerPlanGet AI match score →

Site Reliability Engineer - AI Accelerator Infrastructure - Contract

Santa Clara💼 Contract🗓 2026-07-14 → 2026-07-31

Core

Operate and automate the reliability of AI inference silicon infrastructure across colocation, on-premises labs, and cloud environments.

Role type

Senior Site Reliability Engineer (Infrastructure)

Builds

Colocation server fleets, on-premises GPU clusters, cloud environments, and customer-facing platform services.

Domain

AI Hardware / Silicon Development / High-Performance Computing

Deliverable

production ML models | infrastructure

Required skills

Linux systems administration, Infrastructure as Code (Terraform, Ansible), Kubernetes operations, high-speed interconnects (InfiniBand, RoCE), incident response, Python/Bash scripting, capacity planning

Preferred skills

Customer-facing platform operations, hybrid cloud environments, HPC job schedulers (Slurm, LSF), Go programming, large-scale fleet automation

Technologies

Terraform, Ansible, Kubernetes, Prometheus, Grafana, DataDog, AWS, Azure, GCP, InfiniBand, RoCE, NVLink, Slurm, LSF

Responsibilities

Own reliability and availability of assigned infrastructure domains; perform hands-on server provisioning and hardware troubleshooting; build automation for host lifecycle management and fleet health checks; design monitoring dashboards and alerting rules; triage and resolve incidents with root cause analysis; support customer-facing platform services and deployments.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗