Lead Site Reliability Engineer
Core
Lead technical initiatives for site reliability support services, bridging gaps between software engineering and operations to improve platform availability, resiliency, and performance.
Role type
Senior IC Site Reliability Engineer (Technical Lead)
Builds
Digital payment products and MDES platform infrastructure
Domain
Payments / Distributed Systems
Required skills
System design, performance engineering, chaos testing, capacity planning, incident management, automation, CI/CD, observability, mentoring (via careerplan.io/jobs/R-291891-lead-site-reliability-engineer-at-mastercard)
Preferred skills
Java, Python, Scala, Kubernetes, Helm, performance tuning, root cause analysis
Technologies
Docker, Kubernetes, Dynatrace, Splunk, Grafana, Prometheus, Jenkins, Bamboo, Concourse, BitBucket, Maven, Gatling, Blazemeter
Responsibilities
Lead technical initiatives for site reliability support services before go-live; Increase, maintain, and communicate service metrics post-launch; Scale systems sustainably through automation; Review production incidents to prevent recurrence; Create and maintain technology roadmaps; Build and maintain robust dashboards reflecting system health; Coach and mentor technical talent.
Seniority
Senior, hands-on IC with leadership responsibilities