Staff Site Reliability Engineer
Core
Lead EarnIn's reliability strategy by embedding AI-first workflows to detect, investigate, and prevent production issues across critical financial services.
Role type
Staff Site Reliability Engineer (AI-first operating model)
Builds
AI-assisted incident response systems, reliability scorecards, automated runbooks, and resilient AWS infrastructure
Domain
Fintech / Earned Wage Access / Cloud Infrastructure (AWS)
Deliverable
production ML models | infrastructure
Required skills
SRE strategy, AI/LLM integration for operations, incident command, capacity planning, observability design, mentorship, cross-team influence
Preferred skills
Experience with Datadog, CloudWatch, incident.io, Slack, EKS, Kafka, DynamoDB, RDS, SQS
Technologies
AWS, EKS, Kafka, DynamoDB, RDS, SQS, Datadog, CloudWatch, incident.io, Slack, LLMs
Responsibilities
Define reliability standards (SLIs, SLOs, error budgets); Overhaul incident lifecycle with AI-assisted triage and root-cause analysis; Build AI agents for alert correlation and runbook automation; Coach engineers on reliability practices and AI workflows; Architect for resilience and graceful degradation in AWS environments
Seniority
Staff, hands-on IC with strategic leadership