Infrastructure Site Reliability Engineer
Core
Run and evolve a globally scaled, performance-intensive infrastructure stack supporting AI/HPC workloads, ensuring 24/7 stability and security.
Role type
Senior Infrastructure Site Reliability Engineer (IC)
Builds
Resilient, scalable bare-metal and virtualized infrastructure for AI and HPC systems
Domain
AI Infrastructure / High-Performance Computing (HPC)
Deliverable
production ML models | infrastructure
Required skills
Linux administration (Ubuntu), system tuning, disk I/O optimization, hardware-level performance tweaks, Out of Band management (IPMI, Redfish), networking fundamentals, infrastructure scripting (Bash, Python, Ansible), observability (Prometheus, Grafana), orchestration platforms (Kubernetes, MAAS, Tinkerbell), ITSM frameworks
Preferred skills
HPC workloads knowledge, GPU-based infrastructure, InfiniBand networks, HPC performance tuning
Technologies
Ubuntu, IPMI, Redfish, Prometheus, Grafana, Kubernetes, MAAS, Tinkerbell, Bash, Python, Ansible, InfiniBand
Responsibilities
Deploy and operate resilient infrastructure for AI/HPC; optimize Linux system configuration and hardware performance; manage bare-metal infrastructure via IPMI/Redfish; build automation scripts and IaC; maintain observability stack; perform incident postmortems and root cause analysis; mentor junior engineers
Seniority
Senior, hands-on IC with mentorship responsibilities