CareerPlanGet AI match score →

Senior Devops Engineer Gcp Required 2 Years Of Experience

💼 Full-time🗓 2026-07-27

Core

Build, operate, and optimize scalable, secure, and observable cloud infrastructure for AI-powered services on Google Cloud Platform.

Role type

Senior DevOps Engineer (GCP)

Builds

Production cloud infrastructure, CI/CD pipelines, and AI service integrations (Gemini API via Vertex AI) for engineering teams.

Domain

Cloud Infrastructure (GCP), AI/ML Services, DevOps

Deliverable

production ML models | infrastructure

Required skills

Google Cloud Platform (GCP), Linux system administration, Docker, CI/CD pipelines, cloud networking, IAM, monitoring and logging stacks (Prometheus, Grafana, Loki), troubleshooting infrastructure and networking issues.

Preferred skills

Cost optimization strategies in cloud environments.

Technologies

Google Cloud Platform, Compute Engine, VPC, Load Balancers, Cloud Storage, Cloud Run, Vertex AI, Gemini API, Docker, Prometheus, Grafana, Loki, Ubuntu, CentOS, Nginx, Apache, Cloudflare.

Responsibilities

Manage and support GCP services including Compute Engine, VPCs, Load Balancers, and Cloud Run; Support and integrate Gemini API via Vertex AI into production environments; Apply GCP cost optimization practices including resource sizing and usage monitoring; Build, deploy, and manage Docker images and containers; Design, maintain, and improve CI/CD pipelines; Configure and maintain networking, security, and IAM roles; Set up and maintain observability tooling including Prometheus, Grafana, and Loki; Configure and maintain web and proxy servers (Apache, Nginx).

Seniority

Senior, hands-on IC

Rewrite
## About the Role We are seeking a DevOps Engineer with strong hands-on experience in Google Cloud Platform (GCP) to help build, operate, and optimize scalable, secure, and observable cloud infrastructure. The role involves managing compute, networking, security, CI/CD pipelines, containerized workloads, and supporting AI-powered services. You will work closely with engineering teams to ensure reliability, performance, and cost efficiency on our AI products. ## Required Skills & Qualifications - Experience supporting AI/ML services on GCP. - Experience working in production environments with high availability requirements. - Hands-on experience with Google Cloud Platform (GCP). - Strong Linux system administration skills. - Experience with Docker and containerized workloads. - Working knowledge of CI/CD pipelines and automation. - Understanding of cloud networking, security, and IAM. - Practical experience with monitoring and logging stacks (Prometheus, Grafana, Loki). - Ability to troubleshoot infrastructure, networking, and application-level issues. ## Key Responsibilities ### Cloud Infrastructure (Google Cloud Platform) - Manage and support GCP services including: - Compute Engine (Virtual Machines) - VPCs, Subnets, Routes - Load Balancers (HTTP(S), TCP/UDP) - Cloud Storage - Cloud Run - Assist with provisioning, configuration, and maintenance of cloud resources. ### Support and integrate Gemini API via Vertex AI into production environments - Environment setup and access configuration - IAM and service account management for Vertex AI - Monitoring usage, performance, and costs related to Gemini API ### Apply basic GCP cost optimization practices - Resource sizing and cleanup - Usage monitoring and billing awareness ### Operating Systems & Runtime - Strong understanding of Linux environments: - Ubuntu - CentOS - Troubleshoot system-level and application-level issues. ### Containers & CI/CD - Build, deploy, and manage Docker images and containers. - Design, maintain, and improve CI/CD pipelines. - Write and maintain shell scripts for automation and operational tasks. ### Networking & Security - Solid understanding of core networking concepts: - DNS - HTTP / HTTPS - SSL / TLS - NAT - Internal vs External IP addressing - Configure and maintain: - Firewall rules - Network policies - Manage IAM roles and permissions following least-privilege principles. - Work with Cloudflare for DNS, security, and traffic management. - Diagnose and resolve issues related to HTTP status codes and errors (4xx, 5xx, etc.). ### Monitoring, Logging & Observability - Set up and maintain observability tooling, including: - Prometheus for metrics collection - Grafana for dashboards and visualization - Loki for log aggregation - Configure alerts and assist in incident response. - Analyze logs and metrics to identify, troubleshoot, and resolve performance and reliability issues. ### Web & Proxy Servers - Configure and maintain: - Apache - Nginx - Support reverse proxying, load balancing, and SSL termination. ## Nice to Have - Familiarity with cost optimization strategies in cloud environments.
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗