Site Reliability Engineer
Core
First line of defense for production incidents, monitoring API health, and building diagnostic tools for the API Operations team.
Role type
Site Reliability Engineer (SRE)
Builds
Observability, logging, alerting best practices, and diagnostic tools for API services.
Domain
Payments technology / Cloud infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Incident management, root cause analysis, Python, shell scripting, JavaScript, cloud infrastructure (GCP), observability tools, API integrations
Preferred skills
CI/CD tools, monitoring automation
Technologies
Apigee, GCP, Python, JavaScript, shell scripting
Responsibilities
Triage and resolve production incidents, monitor system health and performance of APIs, implement observability and alerting best practices, build and maintain diagnostic tools and runbooks, partner with governance teams for incident documentation and retrospectives
Seniority
Mid-level, hands-on IC