CareerPlanGet AI match score →
💼 Full-time🗓 2026-06-25

Core

Operate and improve the reliability, observability, and infrastructure of a trading-critical brokerage platform, with a specific focus on deepening PostgreSQL database ownership.

Role type

Senior Site Reliability Engineer (PostgreSQL focus)

Builds

Cloud infrastructure, Kubernetes workloads, observability stack, and PostgreSQL data layer

Domain

Fintech / Brokerage / Cloud Infrastructure

Deliverable

production ML models | product features | dashboards & analysis | infrastructure

Required skills

Kubernetes operations, GitOps, PostgreSQL (performance tuning, online migrations, HA/DR), Linux, observability stack, incident response, Go or Python

Preferred skills

Large-scale OLTP PostgreSQL clusters, typed SQL access layers (pgx, gorm, sqlc), messaging systems (RabbitMQ, Kafka, Redpanda), security/compliance (SOC 2)

Technologies

Kubernetes, GitOps, PostgreSQL, Go, Python, Linux, RabbitMQ, Kafka, Redpanda

Responsibilities

Operate production services including oncall, incident response, and postmortems; Define and refine SLIs/SLOs and error budgets; Strengthen observability across metrics, logs, and traces; Ship infrastructure via code in a GitOps workflow; Manage PostgreSQL performance, schema migrations, and HA/DR; Mentor engineers on reliability and database fundamentals

Seniority

Senior, hands-on IC

Rewrite
## About the Role As a Site Reliability Engineer at Alpaca, you'll help keep our brokerage platform reliable, observable, and operable as we grow - working across our cloud infrastructure, Kubernetes platform, observability stack, messaging layer, and data layer. We're especially interested in candidates with strong PostgreSQL fundamentals who'd like to grow into deeper ownership of our database reliability posture: PostgreSQL sits on the trading-critical path, and we want this person to spend a meaningful share of their time leveling it up while still being a well-rounded SRE the rest of the week. ## Responsibilities - Operate production day-to-day - oncall, incident response, postmortems, and the follow-ups that actually close the loop. - Own reliability practice - define and refine SLIs/SLOs and error budgets, and help product teams live within them. - Strengthen our observability across metrics, logs, traces, and alerting. - Ship infrastructure through code in a GitOps workflow - cloud resources and Kubernetes workloads alike. - Look after PostgreSQL: performance tuning, schema and migration review, online migrations on large tables, HA/DR, and CDC pipelines. - Mentor engineers on reliability and database fundamentals through code review, design review, and pairing. ## Requirements - 4+ years in SRE, DevOps, Platform/Infrastructure, or backend engineering with significant production operations ownership. - Hands-on experience operating production services on Kubernetes, and shipping infrastructure as code in a GitOps workflow. - Solid working knowledge of PostgreSQL in production — query plans, pgstat*, indexing and schema trade-offs, and what a safe online migration looks like on a non-trivial table. - Cloud networking fundamentals (VPCs, routing, L4/L7 load balancing, DNS, TLS) and comfort debugging cross-service connectivity. - Comfortable with a modern observability stack and proficient with Linux at the operator level. - Practiced in incident response - calm under pressure, structured debugging, postmortems that drive change. - At least working proficiency in Go or Python, plus strong written and verbal communication. - Genuine interest in databases and in growing your PostgreSQL/DBA expertise. ## Nice to Have - Deeper PostgreSQL experience: large clusters at OLTP load, online migrations on big tables, HA/DR ownership, connection pooling at scale, or change-data-capture pipelines. - Experience with typed SQL access layers in Go (e.g. pgx, gorm, sqlc). - Production experience with messaging systems at scale (e.g. RabbitMQ, Kafka, Redpanda). - Security & compliance experience in a regulated environment (SOC 2, secrets management, audit logging). - Familiarity with trading, brokerage, or other regulated fintech domains.
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗