CareerPlanSign in

Software Engineer, Site Reliability

San Francisco💼 Full-time🗓 2026-02-23 → 2026-09-25

Core

Own and operate Kubernetes infrastructure at scale, ensuring reliability and availability of customer-facing systems from clusters to networking layers.

Role type

Senior Site Reliability Engineer (Infrastructure)

Builds

Kubernetes clusters, CI/CD pipelines, monitoring dashboards, and incident response processes

Domain

Generative AI infrastructure, cloud-native systems, and high-performance computing

Deliverable

production ML models | infrastructure

Required skills

Kubernetes cluster management, Linux networking, CI/CD pipeline design, Python, Go or Bash, SLO definition, chaos engineering

Preferred skills

GPU/AI workload management, eBPF/XDP, security tooling, distributed storage systems

Technologies

Kubernetes, Terraform, Ansible, Prometheus, Grafana, Loki, Thanos, VictoriaMetrics, Datadog, FluxCD, ArgoCD, Calico, Cilium, MetalLB, Ceph, Longhorn

Responsibilities

Manage Kubernetes cluster lifecycle, upgrades, and multi-tenant isolation; Build and maintain CI/CD pipelines; Leverage AI to automate incident analysis and resolution; Define and enforce SLOs; Manage networking, load balancing, and service mesh configurations; Drive reliability improvements through automation and chaos engineering

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.