Senior Site Reliability Engineer, Observability
Core
Building instrumentation, designing alert configurations, authoring Terraform, and troubleshooting production systems for enterprise treasury customers while coaching product teams on operational maturity.
Role type
Senior Site Reliability Engineer (Observability)
Builds
Observability infrastructure, incident management workflows, and operational maturity for payment systems
Domain
FinTech / Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Terraform, PowerShell, New Relic (APM/Infrastructure/Logs/Synthetics), NRQL, SLO/SLI design, incident management platform configuration, Azure cloud, distributed tracing, structured logging
Preferred skills
Alert noise reduction, chaos engineering, SQL Server monitoring, FinTech compliance (SOC 2/ISO 27001), Python/Bash scripting
Technologies
New Relic, Terraform, Azure, AWS, Incident.IO, OpsGenie, PagerDuty, Azure DevOps, Octopus Deploy, Slack
Responsibilities
Design and implement monitoring, alerting, and dashboards in New Relic; Define and implement SLOs/SLIs and error budgets; Lead alert noise reduction and signal quality engineering; Develop and maintain Terraform infrastructure as code; Administer and configure Incident.IO for alert routing and runbook management; Respond to and debrief on production incidents
Seniority
Senior, hands-on IC