Site Reliability Engineer
Core
Building and scaling cloud infrastructure for a payments-as-a-service platform, ensuring high uptime and reliability for embedded payment solutions.
Role type
Site Reliability Engineer (SRE)
Builds
AWS-based cloud infrastructure, EKS and serverless environments, CI/CD pipelines, and Postgres database infrastructure.
Domain
Fintech / Payments-as-a-service
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Cloud infrastructure (AWS, GCP, Azure), Infrastructure as Code (Terraform, OpenTofu, Terragrunt, CloudFormation), Container orchestration (Kubernetes, ECS), CI/CD pipeline design, Monitoring and observability, Programming (Python, Java, Go, Ruby)
Preferred skills
Startup or high-growth experience, SRE best practices (SLIs, SLOs, error budgets), Cost optimization
Technologies
AWS, Terraform, Kubernetes, GitLab, OpenTelemetry, Prometheus, New Relic, Postgres, Aurora
Responsibilities
Own and scale AWS cloud infrastructure using IaC; Build and operate EKS and serverless environments; Design and maintain CI/CD pipelines; Implement monitoring, alerting, and observability; Automate infrastructure and operational processes; Work with application engineers to improve system performance; Lead incident response and postmortems; Define and roll out SRE best practices; Optimize for cost, security, and compliance; Support and scale Postgres database infrastructure.
Seniority
Mid-level, hands-on IC