Principal Associate SRE
Core
Building a Site Reliability Engineering center in Mexico City to own reliability outcomes independently for payment-critical systems across the Discover Network, Diners Club International, and PULSE.
Role type
Principal Associate SRE (Founding Team)
Builds
Settlement reliability, alert quality, observability dashboards, and automation for millions of daily transactions.
Domain
Financial services / Payments / High-availability infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
SRE/production operations/reliability engineering background, Java/Python/Go, Cloud Native technologies (AWS/Azure/GCP), container orchestration (Docker/Kubernetes), Shell/Bash scripting, Unix/Linux system administration, distributed systems troubleshooting, agentic AI tools (Claude Code)
Preferred skills
Payments/financial services domain experience, Networking concepts (TCP/DNS/TLS), troubleshooting across distributed systems
Technologies
Python, Java, Shell, AWS, Kubernetes, OpenShift, Datadog, Observe, HashiCorp Vault, CI/CD pipelines, agentic AI tools (Claude Code)
Responsibilities
Build and maintain reliability tooling (dashboards, alerts, runbooks, remediation scripts), develop automation solutions to eliminate manual processes, troubleshoot and debug complex production issues across on-prem and cloud, configure and tune monitoring, support incident response and postmortems, manage secrets and certificates, deliver via CI/CD pipelines
Seniority
Senior, hands-on IC (Founding team)