Sr. Site Reliability Engineer
Core
Own the reliability, scalability, and observability of critical financial SaaS applications and infrastructure to ensure seamless, secure, and performant services.
Role type
Senior Site Reliability Engineer (IC)
Builds
Cloud infrastructure, automation tooling, observability architectures, and self-healing systems for fintech SaaS.
Domain
Fintech / Cloud Infrastructure / Observability
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
SLO/SLI design, observability strategy, incident management, cloud architecture (AWS/Azure), infrastructure-as-code, Python scripting, AIOps, chaos engineering, distributed systems knowledge
Preferred skills
Kubernetes/container orchestration, observability as code, network security, database optimization, regulated industry experience (SOC 2, PCI-DSS)
Technologies
AWS, Azure, Terraform, CloudFormation, Ansible, Prometheus, Grafana, ELK, Datadog, New Relic, Python, PowerShell, bash, Gremlin
Responsibilities
Design and maintain SLOs/SLIs; lead observability strategy (monitoring, logging, tracing); build runbooks and incident response procedures; architect cloud infrastructure with high availability; develop automation and AIOps capabilities; drive reliability via load testing and chaos engineering; partner with backend teams on system design; write production-grade Python tooling; champion security and compliance in infrastructure
Seniority
Senior, hands-on IC