Sr. Manager SRE (Individual Contributor)
Core
Senior IC SRE defining reliability vision and driving cross-team execution for three payment networks, pioneering AI-driven automation and observability convergence.
Role type
Senior IC Site Reliability Engineering Manager
Builds
Reliability platform, observability systems, and AI-powered automation for payment settlement networks
Domain
Financial services / Payments / High-availability infrastructure
Deliverable
production ML models | infrastructure
Required skills
SRE/production operations, Java/Python/Go, Cloud Native (AWS/Azure/GCP), Kubernetes/Docker, Unix/Linux administration, Shell/Bash scripting, distributed system debugging, AI/LLM frameworks, observability platform design
Preferred skills
Agentic AI automation development, regulated high-availability domain experience, networking concepts (TCP/DNS/TLS)
Technologies
Python, Java, Shell, AWS, Kubernetes, OpenShift, Datadog, Observe, HashiCorp Vault, Claude Code, LLM frameworks
Responsibilities
Define 12-18 month technical roadmap for GPN SRE, drive reliability transformation (SLOs, error budgets), pioneer AI automation for alert classification and remediation, lead cross-team observability convergence, serve as senior escalation point for complex incidents, architect automation for high-risk processes, mentor engineers on system design and debugging
Seniority
Senior, hands-on IC with strategic leadership