Senior Site Reliability Engineer (Performance and Scalability)
Core
Build platform capacity to absorb campaign-level traffic spikes and enable engineering teams to load- and failure-test their own systems.
Role type
Senior Site Reliability Engineer (Performance and Scalability)
Builds
Scalability foundation, load/failure testing frameworks, observability stack, and hardened infrastructure
Domain
E-commerce / High-scale web platforms
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Capacity planning, autoscaling, caching, queueing, graceful degradation, SLO management, error budget management, observability stack design, Postgres performance tuning, AWS infrastructure, infrastructure-as-code, scripting in Go/TypeScript
Preferred skills
Experience scaling systems through real traffic spikes, designing adopted load/failure testing programs, incident response leadership, blameless postmortems, cross-team enablement
Technologies
AWS, Postgres, TypeScript, Go, PHP/Laravel, observability tooling
Responsibilities
Design capacity planning and autoscaling strategies for large campaign spikes; Establish load and failure testing as standard engineering practice; Own SLOs, error budgets, and observability stack across multiple service languages; Harden Postgres and AWS infrastructure for performance and availability; Lead incident response and drive systemic fixes upstream; Partner with engineering teams to enable self-sufficiency in scaling
Seniority
Senior, hands-on IC