Resiliency Engineer
Core
Design, build, and continuously improve the reliability, availability, and recoverability of technology platforms across on-premises, hybrid, and cloud environments.
Role type
Resiliency Engineer (SRE)
Builds
Automated failover and recovery workflows, self-healing patterns, and disaster recovery test environments
Domain
Banking / Cloud Infrastructure & Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Python, Bash, PowerShell, Terraform, Ansible, Infrastructure-as-Code, High-availability design, Cloud platforms (GCP/Azure), On-premises infrastructure management, Chaos engineering, SLI/SLO definition
Preferred skills
Kubernetes/GKE, CI/CD, Observability tooling (Dynatrace, Prometheus, Grafana, Splunk), ServiceNow, Regulated industry experience, FFIEC/NIST/ISO standards knowledge
Technologies
Python, Bash, PowerShell, Terraform, Ansible, GCP, Azure, Kubernetes, ServiceNow, Dynatrace, Prometheus, Grafana, Splunk, Google FIT, Gremlin
Responsibilities
Design and code automation to replace manual runbooks with orchestrated failover workflows; Develop and maintain infrastructure-as-code to provision and validate recovery environments; Partner with teams to assess architecture for redundancy and meet RTO/RPO targets; Plan and execute disaster recovery tests and fault-injection experiments; Facilitate resiliency reviews and cross-team recovery walkthroughs
Seniority
Mid-Senior, hands-on IC