Lead Site Reliability Engineer
Core
Lead SRE ensuring reliability, scalability, and performance of applications and foundational AI platforms for Mastercard's global operations.
Role type
Lead Site Reliability Engineer (SRE)
Builds
Enterprise AI platforms, tooling, operational practices, and cloud infrastructure
Domain
Payments / AI / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Unix, Shell Scripting, SQL, Python, Apache Nifi, Splunk, Dynatrace, Jenkins, GIT, CI/CD pipeline management, system design, capacity planning, incident response, automation, mentoring
Preferred skills
C, C++, Java, Go, Perl, Ruby, production AI/ML/data platforms, large-scale distributed systems
Technologies
Apache Nifi, Splunk, Dynatrace, Jenkins, GIT, Maven, Artifactory, Chef
Responsibilities
Engage in and improve the whole lifecycle of services from inception to refinement; Analyse ITSM activities and provide feedback on operational gaps; Support services pre-launch via system design consulting and capacity planning; Maintain live services by measuring availability, latency, and system health; Scale systems sustainably through automation; Support application CI/CD pipeline validation and operational gating; Practice sustainable incident response and blameless post-mortems; Mentor junior resources
Seniority
Senior, hands-on IC with mentorship responsibilities