The Digital Site Reliability Engineer (SRE) - GCP Cloud Adoption Engineer
Core
Facilitate the migration, adoption, and optimization of Google Cloud Platform (GCP) services, ensuring digital services are robust, efficient, and aligned with business objectives.
Role type
Senior IC Site Reliability Engineer (GCP Cloud Adoption)
Builds
Cloud infrastructure, automated deployment pipelines, and migration strategies for GCP
Domain
Cloud Computing / Site Reliability Engineering
Deliverable
production ML models | product features | infrastructure
Required skills
GCP services, Infrastructure as Code (Terraform, Deployment Manager), monitoring and logging (Prometheus, Grafana), CI/CD pipeline design, incident management, SRE principles (error budgets, SLIs/SLOs), cloud security and compliance
Preferred skills
Kubernetes, Docker, container orchestration, workload migration from on-premises or other clouds, agile methodologies, project management tools
Technologies
Google Cloud Platform, Terraform, Deployment Manager, Prometheus, Grafana, Kubernetes, Docker
Responsibilities
Develop and execute GCP adoption strategies including migration planning and architecture design; Apply SRE principles to ensure service reliability, availability, and scalability; Build, maintain, and optimize cloud infrastructure using IaC tools; Design and implement automated deployment pipelines and operational workflows; Lead incident response for cloud-related issues and conduct root cause analysis; Monitor system performance to identify improvements in cost, efficiency, and reliability; Ensure cloud environments adhere to security best practices and compliance requirements; Create technical documentation and mentor team members on GCP adoption and SRE practices