Sr. Staff Site Reliability Engineer
Core
Apply software and systems engineering practices to improve the reliability, resilience, scalability, and operational health of production services for the U.S. financial system.
Role type
Senior Staff Site Reliability Engineer (broad technical leadership)
Builds
Production services, reliability tooling, automation, and observability platforms for financial transactions
Domain
Financial services / Cloud infrastructure / Distributed systems
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Software engineering, distributed systems, automation, observability, public cloud (AWS), Linux/Unix, incident response, CI/CD, Infrastructure as Code, capacity management, error budget management
Preferred skills
Hands-on AWS experience, container orchestration, disaster recovery, reusable automation patterns, performance analysis
Technologies
AWS, Linux/Unix, CI/CD, containers, monitoring tools, logging, tracing
Responsibilities
Define and improve SLIs, SLOs, and error budgets; lead incident response and post-incident learning; reduce operational toil through automation; align teams on sustainable reliability practices; develop technical leaders and establish technical direction across a major domain
Seniority
Senior Staff, broad technical direction and force multiplier