Site Reliability Engineer
Core
Operate as a hands-on site reliability engineer delivering after-hours incident detection, triage, and runbook-based remediation for production cloud-native environments.
Role type
hands-on IC site reliability engineer
Builds
production cloud-native environments for North American customers
Domain
Cloud infrastructure (AWS/GCP) and Kubernetes
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Kubernetes operations, AWS/GCP core services, relational database operational basics, observability platforms, Bash/Python scripting
Preferred skills
CKA/CKAD certification, AWS/GCP associate-level certification, incident response experience, on-call/shift operations experience
Technologies
Datadog, Jira, ServiceNow, Confluence, GitHub Wiki, AWS, GCP, Slack, Microsoft Teams
Responsibilities
Monitor and respond to alerts; execute approved runbooks; investigate and contain availability/performance/scaling issues; engage cloud-provider support; prepare escalation summaries and shift-handoff notes; contribute to runbook improvements
Seniority
Mid-level, hands-on IC