Site Reliability Engineer- Product Reliability Engineering
Core
Maintaining, improving, and scaling the reliability of production systems for Visa's global payments infrastructure.
Role type
Site Reliability Engineer (Product Reliability Engineering)
Builds
Automation tooling, monitoring solutions, and incident response workflows for production applications.
Domain
Payments technology, cloud infrastructure, and distributed systems.
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Incident response and troubleshooting, automation scripting, Linux/Unix systems, networking fundamentals, monitoring and observability, CI/CD pipelines, cloud environments (AWS/GCP/Azure), Git version control.
Preferred skills
AI/ML tooling for operations, infrastructure as code, configuration management.
Responsibilities
Participate in incident response and SWAT calls to restore service, perform ongoing analysis to identify potential issues, build automation to reduce manual work, partner with engineering teams on scalability solutions, contribute to monitoring and alerting improvements, develop operational documentation.
Seniority
Mid-level, hands-on IC