Site Reliability Engineer
Core
Operate and improve the reliability, observability, and infrastructure of a trading-critical brokerage platform, with a specific focus on deepening PostgreSQL database ownership.
Role type
Senior Site Reliability Engineer (PostgreSQL focus)
Builds
Cloud infrastructure, Kubernetes workloads, observability stack, and PostgreSQL data layer
Domain
Fintech / Brokerage / Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Kubernetes operations, GitOps, PostgreSQL (performance tuning, online migrations, HA/DR), Linux, observability stack, incident response, Go or Python
Preferred skills
Large-scale OLTP PostgreSQL clusters, typed SQL access layers (pgx, gorm, sqlc), messaging systems (RabbitMQ, Kafka, Redpanda), security/compliance (SOC 2)
Technologies
Kubernetes, GitOps, PostgreSQL, Go, Python, Linux, RabbitMQ, Kafka, Redpanda
Responsibilities
Operate production services including oncall, incident response, and postmortems; Define and refine SLIs/SLOs and error budgets; Strengthen observability across metrics, logs, and traces; Ship infrastructure via code in a GitOps workflow; Manage PostgreSQL performance, schema migrations, and HA/DR; Mentor engineers on reliability and database fundamentals
Seniority
Senior, hands-on IC