CareerPlanGet AI match score →
🌐 Remote💼 Full-time🗓 2026-06-25

Core

Own and drive the entire DevOps function for a high-traffic, production-grade platform serving enterprise customers, focusing on infrastructure reliability, scalability, and delivery velocity.

Role type

Senior DevOps Engineer (Hands-on IC with leadership scope)

Builds

Cloud infrastructure, Kubernetes clusters, CI/CD pipelines, observability stacks, and database systems for enterprise customers

Domain

Cloud Infrastructure & DevOps

Deliverable

infrastructure

Required skills

Google Cloud Platform (GCP) expertise, Kubernetes at scale, Terraform, GitOps (ArgoCD), CI/CD (GitHub Actions), Observability (Grafana/Prometheus/Loki/Tempo), Database operations (PostgreSQL/Redis), Security (Zero-trust/Secrets management), Incident response, Performance engineering

Preferred skills

Cloudflare, Teleport, k6, Multi-cloud environments (AWS)

Technologies

GCP, AWS, Kubernetes, GKE, Terraform, ArgoCD, GitHub Actions, Grafana, Loki, Tempo, Prometheus, Alloy, Cloud SQL, Memorystore, CloudNativePG, MongoDB Atlas, Cloudflare, Teleport, k6

Responsibilities

Manage and evolve GKE clusters running production workloads; Maintain and extend Terraform codebase across multiple projects and accounts; Own GitOps delivery pipeline with automated image updates and self-healing sync; Operate CI/CD systems for build, test, and deployment; Design and operate monitoring stacks for metrics, logs, and tracing; Manage database infrastructure including HA configurations and performance tuning; Implement security posture with secrets management and zero-trust access; Monitor uptime and handle incident response; Execute performance engineering with load testing

Seniority

Senior, hands-on IC

Rewrite
## About the Role We're looking for a Senior DevOps Engineer who will own and drive the entire DevOps function of our platform. This isn't a "maintain the pipeline" role — this is a hands-on leadership position where you take full ownership of infrastructure, reliability, and delivery velocity across the organization. You'll be the person the engineering org leans on when it comes to uptime, scalability, and shipping fast without breaking things. We operate a high-traffic, production-grade platform serving enterprise customers. Our infrastructure is complex, distributed, and built to scale. We need someone who has been through the fire, someone who has operated large-scale systems under real pressure and knows what it takes to keep them resilient, observable, and fast. If you thrive in high-velocity environments, have strong opinions on how infrastructure should be built, and move with urgency, we want to talk to you. ## What You'll Own - Full ownership of our cloud infrastructure across GCP (primary) and AWS, including compute, networking, storage, databases, and security - Kubernetes at scale — managing and evolving our GKE clusters (regional and zonal) running production workloads on high-performance node pools (c3d-highcpu series) - Infrastructure as Code — maintaining and extending our Terraform codebase across multiple GCP projects, AWS accounts, and Cloudflare zones - GitOps delivery pipeline — owning our ArgoCD-driven deployment workflow with automated image updates, self-healing sync, and Slack-integrated deployment notifications - CI/CD systems — GitHub Actions pipelines for build, test, lint, image publishing, and Helm chart distribution to Google Artifact Registry - Observability stack — full Grafana ecosystem including Grafana, Loki (log aggregation), Tempo (distributed tracing), Prometheus (metrics), and Alloy (telemetry pipelines), all backed by GCS and Minio - Database infrastructure — Cloud SQL PostgreSQL 17 (Enterprise HA with pgvector, pgcron, pgstat_statements), Memorystore Redis 7.2 (Standard HA), CloudNativePG operator, and MongoDB Atlas - Networking and edge — Cloudflare DNS/CDN/WAF across multiple domains, GCP VPCs with private subnets, Cloud NAT, private service access, and global load balancing - Security posture — cert-manager with Let's Encrypt, External Secrets Operator syncing from GCP Secret Manager, Cloud KMS key management, Workload Identity Federation, and Teleport for zero-trust access to clusters, databases, and internal applications - Uptime and incident response — Better Stack and UptimeRobot monitoring across all production and development endpoints with escalation policies and status pages - Performance engineering — k6-based load testing infrastructure with custom test scenarios for availability and throughput validation ## What We're Looking For ### Must-Haves - 7+ years of hands-on DevOps/SRE/Infrastructure experience, with at least 3 years operating at senior level in high-scale environments - Deep expertise with Google Cloud Platform — GKE, Cloud SQL, Memorystore, VPC networking, IAM, Workload Identity, Cloud KMS, Artifact Registry, Cloud NAT, and Cloud DNS - Production Kubernetes mastery — you've run, scaled, debugged, and recovered Kubernetes clusters handling real enterprise traffic. Helm chart authoring and management is second nature - Terraform proficiency — you write clean, modular Terraform with remote state, multiple providers (GCP, AWS, Cloudflare, Helm, Kubernetes), and can manage complex multi-environment configurations - Strong GitOps experience — ArgoCD (or equivalent) for continuous delivery with automated sync, image update strategies, and environment promotion - CI/CD pipeline ownership — GitHub Actions (or equivalent) for building, testing, publishing containers, and managing release automation - Observability expertise — designing and operating monitoring stacks (Grafana, Prometheus, Loki, Tempo or similar). You know how to instrument systems, build actionable dashboards, set meaningful alerts, and trace issues across distributed services - Database operations — managing PostgreSQL and Redis in production at scale, including HA configurations, backups, point-in-time recovery, and performance tuning - Security-first mindset — secrets management, zero-trust access patterns, certificate automation, IAM least-privilege, and encryption at rest/in transit - Enterprise resilience — you've designed and operated systems where downtime means real business impact. You understand HA architectures, disaster recovery, failover strategies, and capacity planning - Bias for speed — you ship fast, iterate constantly, and don't let perfect be the enemy of done. You know when to move quickly and when to be careful ### Strong Preferences - Experience with Cloudflare (DNS, WAF, Workers, redirect rulesets) - Experience with Teleport for secure infrastructure access and audit - Experience with k6 or similar tools for performance and load testing - Experience with PostgreSQL in production - Familiarity with release automation tooling and scripting (shell, Python) - Experience operating in multi-cloud environments (GCP + AWS) ## Who You Are - You take ownership. When something is your responsibility, you don't wait to be told what to do. You see the problem, you fix it, you improve the system so it doesn't happen again. - You move fast. You understand that velocity matters. You ship infrastructure changes with confidence because you've built the guardrails — not because you skip them. - You've seen scale. You've operated systems handling serious traffic. You know what breaks at scale and how to prevent it. You've been paged at 3 AM and you've built the systems that stop those pages from happening again. - You're a team player. You work closely with product engineering, you unblock developers, you make the platform better for everyone. - You communicate clearly. You can explain complex infrastructure decisions to technical and non-technical audiences.
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗