Principal Associate (Site Reliability Engineering)
Core
Building a Site Reliability Engineering center in Mexico City to ensure reliability of payment-critical systems across the Discover Network, Diners Club International, and PULSE.
Role type
Principal Associate Site Reliability Engineer (founding team)
Builds
Settlement cycles, automated runbooks, observability dashboards, and AI-powered remediation workflows for global financial transaction systems.
Domain
Financial services / Payments / High-availability infrastructure
Required skills
SRE/production operations background, Java/Python/Go, Cloud Native (AWS/Azure/GCP), Kubernetes/Docker, Shell/Bash scripting, Unix/Linux administration, CI/CD pipelines, observability (Datadog/Observe), secret management (HashiCorp Vault), agentic AI tools (Claude Code) (via careerplan.io/jobs/R1000680-principal-associate-site-reliability-engineering-at-capitalone)
Preferred skills
Troubleshooting distributed systems, payments/financial services domain knowledge, Networking concepts (TCP/DNS/TLS), experience with agentic AI tools
Technologies
Python, Java, Shell, AWS, Kubernetes, OpenShift, Datadog, Observe, HashiCorp Vault, CI/CD pipelines, Claude Code
Responsibilities
Build and maintain reliability tooling (dashboards, alerts, runbooks), develop automation solutions to eliminate manual processes, troubleshoot and debug complex production issues across on-prem and cloud, configure and tune monitoring, support incident response and postmortems, leverage AI tools to accelerate engineering output, manage secrets and certificates
Seniority
Principal, hands-on IC with founding team responsibilities