CareerPlanSign in

Software Engineer, Site Reliability

💼 Full-time🗓 2026-03-13 → 2026-09-25

Core

Own and operate Kubernetes infrastructure, CI/CD pipelines, and networking layers for customer-facing generative media systems at scale.

Role type

Senior Site Reliability Engineer (Infrastructure & Networking)

Builds

High-performance inference platforms, orchestration systems, and observability tools for AI-native products

Domain

Generative AI / Cloud Infrastructure

Deliverable

production ML models | infrastructure

Required skills

Kubernetes cluster management, Infrastructure as Code (Terraform, Ansible), Linux networking, CI/CD and GitOps, Python, Go or Bash, SLO definition, Chaos engineering

Preferred skills

GPU and AI/ML workload management, Kernel-based monitoring (eBPF, XDP), Security tooling (Falco, SIEM), Distributed storage systems (Ceph, Longhorn)

Technologies

Kubernetes, Terraform, Ansible, Prometheus, Grafana, Loki, Thanos, VictoriaMetrics, Datadog, FluxCD, ArgoCD, Calico, Cilium, MetalLB, Ceph, Longhorn, eBPF, XDP, Falco

Responsibilities

Manage Kubernetes cluster lifecycle, upgrades, and multi-tenant isolation; Build and maintain CI/CD pipelines; Automate production issue analysis and resolution using AI; Define and enforce SLOs; Manage networking, load balancing, and service mesh configurations; Drive reliability improvements through automation and chaos engineering

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.