CareerPlanSign in

Site Reliability Engineer

🌐 Remote💼 Full-time🗓 2026-10-01

Core

Own uptime, performance, and observability for a petabyte-scale market data platform serving finance and fintech institutions.

Role type

mid-senior IC Site Reliability Engineer

Builds

high-availability API and platform services for financial data providers

Domain

Financial technology / Market data infrastructure

Deliverable

production ML models | infrastructure

Required skills

Linux debugging and profiling, Python performance optimization, containerization and high-availability deployment, observability tooling, incident response, CI/CD workflow improvement

Preferred skills

Infrastructure-as-code, database query optimization, load testing and capacity planning, alerting best practices

Technologies

Prometheus, OpenTelemetry, VictoriaMetrics, Jaeger, Logstash, Loki, Vector, Docker, Podman, Docker Compose, Docker Swarm, Kubernetes, k3s, Terraform, Ansible, strace, perf, eBPF, ss, gdb

Responsibilities

Own uptime, SLAs, and SLOs across API and platform services; Build and maintain observability across logging, metrics, and tracing; Design and run high-availability deployment and containerization strategies; Profile and optimize Python applications for throughput, latency, and cost; Debug production issues down to the OS level; Improve deployment and CI/CD workflows; Join the on-call rotation, lead incident response, and run post-incident reviews

Seniority

Mid-level to Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 912,000+ jobs from 20+ sources.