CareerPlanSign in
Onsite or remote • Glendale+13💼 Full-time🗓 2026-06-25

Core

Own and evolve the systems that keep infrastructure fast, reliable, and developer-friendly, leading observability, CI/CD, and GitOps initiatives.

Role type

Senior Site Reliability Engineer / Platform Engineer

Builds

Production-grade observability stacks, automated CI/CD pipelines, and scalable Kubernetes environments.

Domain

Cloud Infrastructure, DevOps, Data Engineering

Deliverable

production ML models | infrastructure

Required skills

Observability stack management (Prometheus, Thanos, VictoriaMetrics, Grafana, PromQL), Kubernetes administration (EKS), GitOps implementation (ArgoCD, Flux), CI/CD pipeline design, Infrastructure as Code (Terraform, Pulumi), Scripting (Python, Bash, Go), Stateful service management (MySQL, ClickHouse, Kafka), Security best practices in CI/CD.

Preferred skills

Experience with data lake architecture, Real-time data systems, Cost optimization strategies.

Technologies

EKS, FastAPI, MySQL, ClickHouse, Kafka, Prometheus, Thanos, VictoriaMetrics, Grafana, ArgoCD, Flux, Terraform, Pulumi, Python, Bash, Go.

Responsibilities

Own and evolve the observability stack from data collection to dashboarding; Implement GitOps workflows to automate application lifecycles; Maintain and scale EKS clusters for high availability and cost efficiency; Collaborate on data lake architecture and tooling; Define and enforce security best practices within CI/CD and AWS environments; Continuously improve monitoring, alerting, and incident response playbooks.

Seniority

Senior, hands-on IC

Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.