Senior Site Reliability Engineer – Compute Platforms
Core
Design, implement, and support Kubernetes on baremetal and hypervisor platforms in a private cloud environment, focusing on enterprise compute architecture and automation.
Role type
Senior Site Reliability Engineer (Compute Platforms)
Builds
Enterprise compute and hypervisor environments, Bare Metal as a Service (BMaaS), and production-grade Kubernetes clusters on bare metal.
Domain
Cloud Infrastructure / Private Cloud / Compute Hardware
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Kubernetes, OpenStack, KVM hypervisors, Linux (Ubuntu), Infrastructure as Code (Terraform, Ansible), Python, Bash, Redfish APIs, PXE boot, CI/CD pipelines, hardware lifecycle management
Preferred skills
CNCF Certified Kubernetes Administrator (CKA), Certified Kubernetes Security Specialist (CKS), ITIL Foundation/advanced, OpenStack administration
Technologies
Kubernetes, OpenStack, Harvester, Ubuntu, Terraform, Ansible, Git, ArgoCD, Redfish
Responsibilities
Lead architecture and design of enterprise compute and hypervisor platform solutions; Define standards and automation frameworks for bare metal provisioning; Design and implement Bare Metal as a Service (BMaaS) capabilities; Architect and design Kubernetes platforms on bare metal with QoS and Affinity; Architect and validate automated deployments of operating systems and hypervisors; Perform deep troubleshooting across storage, Kubernetes, hypervisors, networking, and Linux systems.
Seniority
Senior, hands-on IC