Site Reliability Engineer
Core
Manage, maintain, and improve the global Tyk Cloud platform, serving as the first line of incident management for clients.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Global multi-region and multi-cloud Tyk Cloud platform
Domain
Cloud Infrastructure & API Management
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Kubernetes & containers, AWS/EKS, Linux, Terraform/IaC, Helm, Go, MongoDB, Redis, Prometheus, Grafana, Networking concepts, Logging systems
Preferred skills
GCP, Azure, Bare metal infrastructure, API management, Large scale distributed storage, Rancher, CKA/CKAD/CKS, Go development
Technologies
Kubernetes, AWS, EKS, MongoDB, Redis, Prometheus, Grafana, Thanos, Terraform, Helm, Go
Responsibilities
Maintain global Tyk Cloud within defined SLAs, identify and solve reliability issues, introduce new metrics and build dashboards, participate in on-call rotation, expand multi-region/multi-cloud reach, document operational knowledge, conduct post-incident analysis, automate common tasks, drive operational efficiency and cost reduction, assist in penetration testing, manage incidents
Seniority
Senior, hands-on IC
