Software Engineer, Site Reliability
Core
Own and operate Kubernetes infrastructure, CI/CD pipelines, and networking layers for customer-facing generative media systems at scale.
Role type
Senior Site Reliability Engineer (Infrastructure & Networking)
Builds
High-performance inference platforms, orchestration systems, and observability tools for AI-native products
Domain
Generative AI / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes cluster management, Infrastructure as Code (Terraform, Ansible), Linux networking, CI/CD and GitOps, Python, Go or Bash, SLO definition, Chaos engineering
Preferred skills
GPU and AI/ML workload management, Kernel-based monitoring (eBPF, XDP), Security tooling (Falco, SIEM), Distributed storage systems (Ceph, Longhorn)
Technologies
Kubernetes, Terraform, Ansible, Prometheus, Grafana, Loki, Thanos, VictoriaMetrics, Datadog, FluxCD, ArgoCD, Calico, Cilium, MetalLB, Ceph, Longhorn, eBPF, XDP, Falco
Responsibilities
Manage Kubernetes cluster lifecycle, upgrades, and multi-tenant isolation; Build and maintain CI/CD pipelines; Automate production issue analysis and resolution using AI; Define and enforce SLOs; Manage networking, load balancing, and service mesh configurations; Drive reliability improvements through automation and chaos engineering
Seniority
Senior, hands-on IC
