Senior Infrastructure/Site Reliability Engineer On-Call
Core
Senior on-call SRE responsible for incident response and stability of enterprise AI infrastructure (Kubernetes, Ceph, bare metal).
Role type
Senior hands-on IC Site Reliability Engineer
Builds
High-availability AI workloads for enterprise clients
Domain
AI Infrastructure / Distributed Systems
Deliverable
production ML models
Required skills
Kubernetes cluster operations, Ceph storage, bare metal troubleshooting, CNI networking (Cilium/Calico), Linux systems administration, etcd cluster management, GPU infrastructure (NVIDIA operator), infrastructure-as-code (Ansible/Kubespray)
Preferred skills
AI inference/training infrastructure experience
Technologies
Kubernetes, Ceph, Cilium, Calico, etcd, Ansible, Kubespray, NVIDIA Kubernetes operator, IPMI
Responsibilities
Respond to and resolve production incidents across client infrastructure, troubleshoot complex distributed systems problems, handle escalations requiring deep expertise in storage and networking, communicate directly with enterprise clients during incidents, document incidents and improve runbooks, collaborate on long-term reliability improvements and automation
Seniority
Senior, hands-on IC