Principal Engineer Software
Core
Architecting resilient physical layouts and automating the deployment, scaling, and self-healing of Red Hat OpenShift production clusters to ensure 99.99% availability.
Role type
Senior Staff Data Center & OpenShift Operations Engineer
Builds
High-availability OpenShift 4.x clusters with Zero Single Point of Failure (ZSPoF) architecture
Domain
Cybersecurity infrastructure / Data Center Operations / Kubernetes
Deliverable
production ML models | infrastructure (via careerplan.io/jobs/JR-019075-principal-engineer-software-at-paloaltonetworks)
Required skills
Red Hat OpenShift 4.x administration, high-density GPU server racking and cabling, Infrastructure as Code (Ansible/Pulumi), Python/Bash scripting, Linux (CoreOS/RHEL) kernel tuning, BGP/VLAN/LACP networking, vSphere/KVM virtualization
Preferred skills
DCIM tools (Netbox), Prometheus/Grafana observability, OADP/Velero disaster recovery, NVIDIA DGX hardware experience
Technologies
Red Hat OpenShift, Ansible, Pulumi, Python, Bash, CoreOS, RHEL, NVIDIA DGX, Dell, Cisco, OVN-Kubernetes, vSphere, KVM, Ceph, OpenShift Data Foundation, Prometheus, Grafana, Netbox, ELK, Lok
Responsibilities
Monitor and maintain data center systems for ZSPoF architecture; Implement cluster reliability engineering across power/cooling zones; Design and execute automated failover strategies; Perform routine maintenance and upgrades using GitOps; Resolve deep-stack hardware and software issues; Coordinate with vendors for specialized hardware; Optimize rack density and thermal loads; Integrate hardware health metrics into observability stacks; Rack and stack high-density GPU servers; Perform precision physical installation and replacement of critical components
Seniority
Senior Staff, hands-on IC