Sr. Site Reliability Engineer
Core
Incident command role directing application and infrastructure teams during critical incidents, authorizing rollbacks and failovers, and driving site resiliency projects.
Role type
Senior Site Reliability Engineer (Incident Command)
Builds
PayPal's core payment platforms (Venmo, Xoom, Zettle, Braintree) and global financial infrastructure
Domain
Financial Services / Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Incident command, Infrastructure as Code (Terraform, CloudFormation), Kubernetes, CI/CD, disaster recovery planning, cloud architecture (AWS, GCP, Azure), cross-functional leadership
Preferred skills
Executive stakeholder management, blameless postmortem facilitation, tooling development for incident response
Technologies
Terraform, CloudFormation, Kubernetes, AWS, GCP, Azure
Responsibilities
Act as incident commander with final decision authority; Direct application and infrastructure teams during incidents; Review Infrastructure as Code changes for reliability risks; Conduct regular disaster recovery drills; Lead blameless postmortem sessions; Drive site resilience projects to enhance system reliability
Seniority
Senior, hands-on IC with strategic decision-making authority