Senior Staff SRE – Compute Platform
Core
Build and operate reliable, scalable compute platforms supporting global engineering workloads, spanning Kubernetes, KubeVirt, bare-metal infrastructure, and AI-enabled operations.
Role type
Senior Staff Site Reliability Engineer (Compute Platform)
Builds
Large-scale Kubernetes, KubeVirt, Linux, container, and bare-metal compute platforms
Domain
Cloud Infrastructure / HPC / AI Compute
Deliverable
production ML models | infrastructure
Required skills
Kubernetes administration, KubeVirt, bare-metal provisioning, Infrastructure as Code, Python, Go, SLO/SLI definition, incident management, distributed systems expertise
Preferred skills
HPC/AI/GPU-accelerated infrastructure operations, VMware vSphere, Red Hat OpenShift, KVM, Firecracker, OpenStack, Nutanix AHV, generative AI for operations
Technologies
Kubernetes, KubeVirt, Docker, Terraform, Ansible, Chef, Puppet, OpenTelemetry, Prometheus, Grafana, ELK Stack, Splunk, Linux
Responsibilities
Lead bare-metal provisioning and lifecycle management in data centers; develop automation and self-service capabilities; define and operate SLOs, SLIs, and error budgets; lead complex incident investigations and postmortems; partner with cross-functional teams on global platform initiatives
Seniority
Senior Staff, hands-on IC with strategic scope
