Principal Platform Engineer (High Availability & Disaster Recovery)
Core
Architecting and executing global high availability and disaster recovery strategies for a SaaS scenario planning platform to ensure continuous uptime and business continuity.
Role type
Principal Platform Engineer (High Availability & Disaster Recovery)
Builds
Global HA architectures, automated failover mechanisms, and a foundational Resiliency Engineering framework for a SaaS platform.
Domain
Cloud & On-Premises Infrastructure, SaaS, Disaster Recovery
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
High Availability engineering, Cloud infrastructure (AWS/GCP/Azure), Kubernetes, Terraform/Ansible, Hybrid networking (BGP/DNS/CDN), Database clustering, Real-time replication, Chaos Engineering, Incident management
Preferred skills
SaaS product support, Observability and monitoring best practices
Technologies
AWS, GCP, Azure, Kubernetes, Terraform, Ansible, BGP, DNS, CDNs
Responsibilities
Own the end-to-end design and technical execution of the global High Availability roadmap; Build and implement active-active clustering and global load balancing; Establish the foundational Resiliency Engineering framework; Partner with engineering squads to embed self-healing mechanisms; Bootstrap and define strategy for the Chaos Engineering program; Monitor infrastructure KPIs to mitigate risks and ensure SLA compliance.
Seniority
Principal, hands-on IC with strategic influence