Platform Engineer IV
Core
Architect, build, and evolve fault-tolerant infrastructure and application patterns for enterprise availability and recovery commitments on AWS.
Role type
Senior IC Site Reliability Engineer (SRE) / Cloud Architect
Builds
Multi-region, multi-AZ resilient architectures, reference patterns, automated self-healing systems, and orchestrated failover mechanisms.
Domain
Financial Services / Cloud Infrastructure / Disaster Recovery
Deliverable
production ML models | product features | infrastructure
Required skills
AWS Backup and recovery orchestration, SLO/SLI definition and error budget management, chaos engineering (Chaos Mesh, Gremlin, FIS), Kubernetes/EKS scaling and networking, Infrastructure-as-Code (Terraform), observability (Prometheus, OpenTelemetry, Datadog), technical leadership and mentorship.
Preferred skills
Workflow/streaming platforms (AutoSys, Temporal.io, Kafka).
Technologies
AWS (Backup, Application Recovery Controller, Resiliency Hub, FIS, EKS), Terraform, Prometheus, OpenTelemetry, Datadog, Chaos Mesh, Gremlin, Kafka, AutoSys, Temporal.io.
Responsibilities
Design multi-region architectures and publish reference patterns; lead chaos experiments and game days; automate self-healing and failover to meet RTO/RPO; coordinate incident response and author post-mortems; mentor senior and lead engineers.
Seniority
Senior, hands-on IC with strategic lead responsibilities.