Staff Site Reliability Engineer - Release Engineering
Core
Define and scale reliability practices across product engineering, architecting SLO and error-budget programs to ensure safe, high-velocity production deployments.
Role type
Staff Site Reliability Engineer (Release Engineering)
Builds
Zero-touch deployment systems, progressive rollout frameworks, metric-gated analysis, and automatic rollback infrastructure.
Domain
Fintech / Cloud Infrastructure / Release Engineering
Deliverable
production ML models | product features | infrastructure
Required skills
Backend systems architecture, SRE program design, progressive delivery implementation, observability strategy, incident response leadership, organizational change management, high-stakes production judgment, Go or similar systems languages, Kubernetes, service mesh technologies, Prometheus, ArgoCD
Preferred skills
Designing service maturity models, SLI frameworks, canary rollout systems, automated rollback infrastructure, driving engineering culture without formal authority
Technologies
Go, Kubernetes, service mesh, Prometheus, ArgoCD
Responsibilities
Lead expansion of reliability standards across product engineering; Architect and manage the SLO and error-budget framework; Promote widespread use of progressive delivery and automated safety gates; Guide emerging product teams toward production readiness; Collaborate with Platform and Infrastructure teams to transform production requirements into self-service features; Direct response to critical incidents and ensure post-mortem actions yield permanent improvements.
Seniority
Staff, hands-on technical leadership