Site Reliability Engineer II
Core
Production readiness steward for Mastercard products, ensuring platform stability, health, and operational standards while fostering developer run ownership.
Role type
Senior IC Site Reliability Engineer
Builds
High-availability systems, automated operational tools, and resilient product infrastructure for global commerce.
Domain
Fintech / Global Payments / Cloud Infrastructure
Deliverable
production ML models | product features | infrastructure
Required skills
Observability (metrics, logs, traces), Programming and Scripting (Python, Go, Bash), Linux/Unix Systems Administration, Cloud Platform Management (AWS, Azure, GCP), High Availability and Disaster Recovery Design, CI/CD and Containerization, Incident Response and Troubleshooting, Capacity Planning, IT Service Management (Incident/Problem/Change), Proactive Monitoring and Reliability Engineering.
Preferred skills
None explicitly stated as preferred, only required capabilities listed.
Technologies
Python, Go, Bash, Linux, Unix, AWS, Azure, GCP, CI/CD pipelines, Containers, Kubernetes (implied by orchestration), Monitoring tools.
Responsibilities
Implement observability solutions for incident detection and diagnosis, write scripts to automate operational tasks and build tools, configure and troubleshoot Linux/Unix systems and network components, design and manage cloud infrastructure for scalability and security, design systems for high availability and fault tolerance, apply DevOps principles for faster software delivery, systematically diagnose and resolve technical issues, monitor resource utilization and forecast capacity needs, apply IT service management principles for incident and change management, document operational procedures and share knowledge.
Seniority
Mid-Senior, hands-on IC with occasional guidance in complex scenarios.