CareerPlanGet AI match score →

Staff Site Reliability Engineer, Database

💼 Full-time🗓 2026-06-25

Core

Ensure reliability, scalability, and performance of multi-terabyte scale PostgreSQL clusters and systems for a fintech trading platform.

Role type

Staff Site Reliability Engineer (Database)

Builds

High-availability, high-performance PostgreSQL database infrastructure and low-latency trading systems

Domain

Fintech / Trading / Database Infrastructure

Deliverable

production ML models | product features | infrastructure

Required skills

PostgreSQL (multi-terabyte scale), Go, Prometheus, Linux, distributed tracing, schema design, query optimization, capacity planning, incident management, SLI/SLO/SLA design

Preferred skills

pgx, gorm, sqlc, low-latency system experience

Technologies

PostgreSQL, Go, Prometheus, Linux, pgx, gorm, sqlc

Responsibilities

Triage difficult technical problems and implement solutions; Improve observability stack (monitoring, logging, profiling); Respond to and resolve incidents and conduct post-incident reviews; Collaborate with development teams on reliability and scalability; Monitor system capacity and implement changes for future growth

Seniority

Staff, hands-on IC

Rewrite
## About the Role As a Site Reliability Engineer (SRE) at Alpaca, you will ensure the reliability, scalability, and performance of our systems and services. You will work closely with development, operations and devops teams to build and maintain robust applications, ensuring they run smoothly and efficiently. This role requires a blend of software engineering and operations skills, with a strong ability to troubleshoot technical issues and resolve problems before they impact our users. ## Responsibilities - Triage difficult technical problems and implement solutions - Improve our observability stack (monitoring, logging, profiling) - Incident Management: Respond to and resolve incidents in a timely manner, conducting post-incident reviews to identify and implement improvements. - Collaboration: Work closely with development teams to ensure new features and services are designed with reliability and scalability in mind. - Capacity Planning: Monitor system capacity and performance, making recommendations and implementing changes to handle future growth. ## Requirements - 5+ years of experience in Site Reliability Engineering, Performance Engineering, or similar roles. - 5+ years of experience with multi-terabyte scale PostgreSQL clusters. - Proven track record of managing and maintaining large-scale, high-availability, and high-performance PostgreSQL database. - Experience designing and implementing SLIs, SLOs, and SLAs for internal systems and databases. - Experience with troubleshooting PostgreSQL performance problems and slow queries. - Extensive experience with efficient schema design and efficient query design. - Experience migrating multi-terabyte tables into more efficient schemas. - Proficient with Go. - Proficient with Prometheus. - Proficient with Linux. - Knowledgeable in trading/fintech domains. - Experience with low-latency systems. - Experience with distributed tracing. - Experience scaling PostgreSQL clusters rapidly. - Experience with pgx, gorm, or sqlc.
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗