Director, Infrastructure & Site Reliability Engineering
Core
Lead strategic initiatives to ensure the reliability, scalability, and performance of VMware and Oracle Linux platforms for Mastercard's distributed infrastructure.
Role type
Director, Site Reliability Engineering (Infrastructure)
Builds
VMware clusters, ESXi hosts, Oracle Linux environments, and automated self-healing systems
Domain
Payments / Enterprise Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
VMware (ESXi, clusters), Oracle Linux, Infrastructure-as-Code, automation frameworks, SRE practices, observability tools, high availability, disaster recovery
Preferred skills
Large-scale infrastructure modernization, executive communication, cross-functional collaboration
Technologies
Chef, Ansible, PowerCLI, Python, Jenkins, Prometheus, Grafana, Splunk, Dynatrace
Responsibilities
Define and execute the strategic roadmap for SRE across distributed platforms; Lead modernization efforts including hardware lifecycle management and virtualization upgrades; Build, mentor, and scale a high-impact SRE organization; Oversee the health and performance of VMware clusters and Oracle Linux environments; Architect observability solutions and define Service Level Objectives (SLOs); Drive incident management and root cause analysis for critical infrastructure issues; Collaborate with InfoSec and audit teams to maintain a secure and compliant environment.
Seniority
Director, strategic leadership with hands-on technical expertise