Site Reliability (DevOps) Engineering Manager
Core
Lead a distributed team of Site Reliability Engineers to ensure the reliability, scalability, and performance of Rakuten's global marketing cloud platforms.
Role type
Senior IC Site Reliability Engineering Manager
Builds
Highly available marketing platforms serving millions of customers globally
Domain
Internet services, marketing technology, cloud infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
People management, SRE strategy (SLO/SLI), incident management, observability design, automation, capacity planning, cloud platform expertise, container orchestration, Infrastructure as Code, CI/CD, scripting
Preferred skills
Big data technologies, marketing technology platforms, database administration, chaos engineering, cloud certifications
Technologies
GCP, AWS, Azure, Kubernetes, Docker, Terraform, Ansible, Prometheus, Grafana, Datadog, ELK Stack, Python, Go, Java, PostgreSQL, MySQL, Redis, Couchbase, Chaos Monkey, Litmus, Gremlin
Responsibilities
Lead and mentor a distributed SRE team across multiple time zones; Define and drive SRE strategy including SLO/SLI frameworks and error budgets; Establish incident management processes and blameless post-mortem practices; Collaborate with development teams to embed reliability practices into the SDLC; Design and implement comprehensive observability solutions; Drive automation initiatives to reduce toil; Partner with Architecture and Platform teams on infrastructure decisions; Manage capacity planning and performance optimization for critical marketing platforms; Report on reliability metrics and operational health to leadership.
Seniority
Senior, hands-on IC with people management