Software Engineer, Site Reliability
Core
Own and operate Kubernetes infrastructure at scale, ensuring reliability and availability of customer-facing systems from clusters to networking layers.
Role type
Senior Site Reliability Engineer (Infrastructure)
Builds
Kubernetes clusters, CI/CD pipelines, monitoring dashboards, and incident response processes
Domain
Generative AI infrastructure, cloud-native systems, and high-performance computing
Deliverable
production ML models | infrastructure
Required skills
Kubernetes cluster management, Linux networking, CI/CD pipeline design, Python, Go or Bash, SLO definition, chaos engineering
Preferred skills
GPU/AI workload management, eBPF/XDP, security tooling, distributed storage systems
Technologies
Kubernetes, Terraform, Ansible, Prometheus, Grafana, Loki, Thanos, VictoriaMetrics, Datadog, FluxCD, ArgoCD, Calico, Cilium, MetalLB, Ceph, Longhorn
Responsibilities
Manage Kubernetes cluster lifecycle, upgrades, and multi-tenant isolation; Build and maintain CI/CD pipelines; Leverage AI to automate incident analysis and resolution; Define and enforce SLOs; Manage networking, load balancing, and service mesh configurations; Drive reliability improvements through automation and chaos engineering
Seniority
Senior, hands-on IC
