Site Reliability Engineer Senior
Core
Building and maintaining high-reliability infrastructure for instant B2B payments across Latin America, enabling zero-downtime deployments and automated treasury solutions.
Role type
Senior Site Reliability Engineer (Infrastructure & SRE Strategy)
Builds
Cloud-native payment infrastructure (AWS, Kubernetes) supporting fintechs, PSPs, and banks
Domain
Fintech / B2B Payments / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Cloud infrastructure (AWS), Kubernetes, Infrastructure as Code (Terraform, Pulumi), Microservices architecture, Incident management, Observability (monitoring, logging, alerting), Capacity planning, Automation scripting (Python), AIOps concepts, Internal Developer Platforms (Backstage, Cortex)
Preferred skills
AI-driven auto-remediation tools (n8n, AWS Bedrock), GitOps patterns, Secure-by-design principles
Technologies
AWS, Kubernetes, Terraform, Pulumi, NewRelic, Datadog, CloudWatch, n8n, AWS Bedrock, Backstage, Cortex
Responsibilities
Define infrastructure supporting solution architecture, support all infrastructure (AWS services and k8s clusters) for zero-downtime deployments, troubleshoot application and infrastructure incidents, review code instrumentation and create monitoring dashboards, document and maintain runbooks while automating incident response, perform load/scalability testing and capacity planning, design peak readiness reviews, lead operational state reviews, contribute to incident management and postmortems, socialize SRE culture and mentor teams
Seniority
Senior, hands-on IC with strategic alignment