SRE Architect, AI-Powered Reliability
Core
Define and enforce enterprise-wide SRE standards, operational practices, and architectural guardrails using AI to scale reliability across mission-critical financial systems.
Role type
Principal SRE Architect (AI-Powered Reliability)
Builds
Enterprise reliability frameworks, AI-driven observability and incident management tooling, self-healing systems, and capacity planning models for global payment and technology solutions.
Domain
Financial technology / Payments / High-availability distributed systems
Deliverable
production ML models | infrastructure
Required skills
SRE and distributed systems architecture, AI-powered reliability engineering, observability strategy, incident management frameworks, resilience engineering, capacity planning, cloud cost optimization, technical governance, self-healing system design, chaos engineering
Preferred skills
Fintech domain expertise, building SRE practices from scratch, AIOps platforms, chaos engineering tools, Kubernetes and service mesh, event-driven architectures
Technologies
AWS, GCP, Azure, Kubernetes, Istio, Linkerd, OpenTelemetry, Dynatrace, Honeycomb, Coralogix, Gremlin, LitmusChaos, AWS Fault Injection Simulator
Responsibilities
Define and enforce enterprise SRE best practices and operational standards; Architect and oversee mission-critical systems for reliability and transactional integrity; Establish SLO, SLI, and error budget frameworks; Lead AI-Powered Reliability Engineering strategy and adoption of SRE agents; Drive enterprise observability standards and AI-assisted anomaly detection; Establish enterprise incident management frameworks and integrate AI for intelligent triage; Lead resilience engineering initiatives and chaos engineering programs; Develop enterprise capacity planning strategies and enforce load testing standards; Drive cloud cost optimization and budgeting initiatives; Design and champion blameless postmortem culture with AI acceleration.
Seniority
Principal, strategy & mentorship