Site Reliability Engineer - AI Accelerator Infrastructure - Contract
Core
Operate and automate the reliability of AI inference silicon infrastructure across colocation, on-premises labs, and cloud environments.
Role type
Senior Site Reliability Engineer (Infrastructure)
Builds
Colocation server fleets, on-premises GPU clusters, cloud environments, and customer-facing platform services.
Domain
AI Hardware / Silicon Development / High-Performance Computing
Deliverable
production ML models | infrastructure
Required skills
Linux systems administration, Infrastructure as Code (Terraform, Ansible), Kubernetes operations, high-speed interconnects (InfiniBand, RoCE), incident response, Python/Bash scripting, capacity planning
Preferred skills
Customer-facing platform operations, hybrid cloud environments, HPC job schedulers (Slurm, LSF), Go programming, large-scale fleet automation
Technologies
Terraform, Ansible, Kubernetes, Prometheus, Grafana, DataDog, AWS, Azure, GCP, InfiniBand, RoCE, NVLink, Slurm, LSF
Responsibilities
Own reliability and availability of assigned infrastructure domains; perform hands-on server provisioning and hardware troubleshooting; build automation for host lifecycle management and fleet health checks; design monitoring dashboards and alerting rules; triage and resolve incidents with root cause analysis; support customer-facing platform services and deployments.
Seniority
Senior, hands-on IC