CareerPlanGet AI match score →
Onsite or remote • Glendale+13💼 Full-time🗓 2026-06-25

Core

Own and evolve the systems that keep infrastructure fast, reliable, and developer-friendly, leading observability, CI/CD, and GitOps initiatives.

Role type

Senior Site Reliability Engineer / Platform Engineer

Builds

Production-grade observability stacks, automated CI/CD pipelines, and scalable Kubernetes environments.

Domain

Cloud Infrastructure, DevOps, Data Engineering

Deliverable

production ML models | infrastructure

Required skills

Observability stack management (Prometheus, Thanos, VictoriaMetrics, Grafana, PromQL), Kubernetes administration (EKS), GitOps implementation (ArgoCD, Flux), CI/CD pipeline design, Infrastructure as Code (Terraform, Pulumi), Scripting (Python, Bash, Go), Stateful service management (MySQL, ClickHouse, Kafka), Security best practices in CI/CD.

Preferred skills

Experience with data lake architecture, Real-time data systems, Cost optimization strategies.

Technologies

EKS, FastAPI, MySQL, ClickHouse, Kafka, Prometheus, Thanos, VictoriaMetrics, Grafana, ArgoCD, Flux, Terraform, Pulumi, Python, Bash, Go.

Responsibilities

Own and evolve the observability stack from data collection to dashboarding; Implement GitOps workflows to automate application lifecycles; Maintain and scale EKS clusters for high availability and cost efficiency; Collaborate on data lake architecture and tooling; Define and enforce security best practices within CI/CD and AWS environments; Continuously improve monitoring, alerting, and incident response playbooks.

Seniority

Senior, hands-on IC

Rewrite
## About the Role As a key member of our Platform Engineering team, you'll own and evolve the systems that keep our infrastructure fast, reliable, and developer-friendly. You'll lead our observability, CI/CD, and GitOps initiatives—empowering engineers to ship confidently and our systems to scale intelligently. You'll work in a high-traffic, data-driven environment built on EKS, FastAPI, MySQL/ClickHouse, and Kafka, helping ensure everything from deployments to dashboards runs smoothly and securely. ## What You'll Do - Own and evolve our observability stack, from data collection through long-term retention and dashboarding. - Implement GitOps workflows with ArgoCD (or similar), automating application lifecycles from commit to production. - Collaborate on data lake architecture and tooling ensuring platform infrastructure supports analytics and data engineering workflows. - Maintain and scale EKS clusters, ensuring high availability, cost efficiency, and performance. - Collaborate closely with backend and infrastructure engineers to standardize environments, automate workflows, and reduce deployment friction. - Define and enforce security best practices within CI/CD and AWS environments. - Continuously improve monitoring, alerting, and incident response playbooks. ## Must-Haves - Observability Expert: Deep Prometheus, Thanos or VictoriaMetrics experience, Grafana, PromQL, SLO design, dashboarding and alerting - K8S Expertise: Solid, hands-on experience managing production workloads on K8S or Amazon EKS (including scaling, upgrades, and cost management). - GitOps Champion: Hands-on experience with ArgoCD, Flux, or similar tools managing production deployments. ## Nice-to-Haves - Experience managing stateful services (e.g., MySQL, KeyDB/Redis, ClickHouse, Cassandra). - Familiarity with Terraform or Pulumi for IaC. - Scripting experience (Python, Bash, or Go) for automation and tooling. - Experience supporting data-heavy or real-time systems (Kafka, analytics platforms, etc.). ## Success Looks Like - Seamless, automated deployments across all environments. - Actionable observability: metrics and alerts that detect issues before users do. - Zero-downtime rollouts and predictable releases. - Developers fully empowered to ship code quickly and safely.
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗