CareerPlanSign in

Senior Infrastructure/Site Reliability Engineer On-Call

San Francisco💼 Full-time🗓 2026-09-30 → 2026-10-01

Core

Senior on-call SRE responsible for incident response and stability of enterprise AI infrastructure (Kubernetes, Ceph, bare metal).

Role type

Senior hands-on IC Site Reliability Engineer

Builds

High-availability AI workloads for enterprise clients

Domain

AI Infrastructure / Distributed Systems

Deliverable

production ML models

Required skills

Kubernetes cluster operations, Ceph storage, bare metal troubleshooting, CNI networking (Cilium/Calico), Linux systems administration, etcd cluster management, GPU infrastructure (NVIDIA operator), infrastructure-as-code (Ansible/Kubespray)

Preferred skills

AI inference/training infrastructure experience

Technologies

Kubernetes, Ceph, Cilium, Calico, etcd, Ansible, Kubespray, NVIDIA Kubernetes operator, IPMI

Responsibilities

Respond to and resolve production incidents across client infrastructure, troubleshoot complex distributed systems problems, handle escalations requiring deep expertise in storage and networking, communicate directly with enterprise clients during incidents, document incidents and improve runbooks, collaborate on long-term reliability improvements and automation

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 902,000+ jobs from 20+ sources.