Lead Site Reliability Engineer (Lead SRE) - Incentive Platform Department (INPD)
Core
Define and drive technical direction for stable operation, continuous improvement, and scalability of mission-critical incentive and payment services.
Role type
Lead Site Reliability Engineer (Lead SRE)
Builds
Incentive and payment platforms (Rakuten Point, Rakuten Coupon) supporting the global Rakuten Ecosystem
Domain
Fintech / E-commerce / Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Kubernetes architecture and operations, cloud platform management (AWS/GCP/Azure), observability stack implementation (Prometheus/Grafana/ELK/Datadog), incident command and root cause analysis, CI/CD automation, system performance tuning, capacity planning, UNIX internals, networking protocols (TCP/IP/HTTP), scripting (Shell/Python), Git/GitHub
Preferred skills
GCP environment development (GKE/Cloud Run/BigQuery), web application development, error budget implementation, cross-cultural global team leadership
Technologies
Kubernetes, AWS, GCP, Azure, Prometheus, Grafana, ELK Stack, Datadog, Jenkins, CircleCI, GitLab, Shell, Python, TCP/IP, HTTP, Git, GitHub
Responsibilities
Define and achieve Service Level Objectives (SLOs) and Service Level Agreements (SLAs); Identify and resolve service performance and latency bottlenecks; Act as incident commander during production outages; Architect scalable operational frameworks and automate operational processes; Provide technical guidance and mentorship to SRE team members; Lead cross-functional collaboration with product, infrastructure, and security teams; Participate in 24/7 on-call rotation and refine escalation paths
Seniority
Senior, hands-on IC with leadership responsibilities