Engineer III - Site Reliability
Core
Own production infrastructure spanning multiple clouds and regions, building automation, hardening security, and enabling self-service capabilities for engineering teams.
Role type
Senior Site Reliability Engineer (Platform Engineering)
Builds
Shared infrastructure platforms, CI/CD pipelines, control planes, observability stacks, and disaster recovery systems.
Domain
Cybersecurity, Cloud Infrastructure, Distributed Systems
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Kubernetes, Infrastructure-as-Code (Terraform/Pulumi/Crossplane), GitOps (ArgoCD/Flux), Observability (Prometheus/Grafana/OpenTelemetry), Database Operations, Security Hardening, Multi-cloud Management, AI-driven Automation
Preferred skills
Security platforms/telemetry pipelines, Internal Developer Platforms, Service Mesh (Istio/Linkerd), Workflow Orchestration (Temporal/Argo Workflows), Go proficiency
Technologies
Kubernetes, GitHub Actions, Jenkins, Tekton, Terraform, Pulumi, Crossplane, ArgoCD, Flux, Prometheus, Grafana, Jaeger, OpenTelemetry, Istio, Linkerd, Temporal, Argo Workflows, Go
Responsibilities
Deploy, upgrade, and maintain platform services across multiple clouds and regions; Build and maintain CI/CD pipelines using GitOps workflows; Create APIs and tooling for provisioning and scaling; Track usage, forecast growth, and manage infrastructure costs; Set up metrics, dashboards, and alerts; Join on-call rotation, resolve incidents, and write postmortems; Automate deployments, upgrades, certificate rotations, and failover; Harden security with auth, encryption, and network policies; Build backup strategies and test failover for disaster recovery; Provide templates and patterns to help engineering teams use platforms reliably.
Seniority
Senior, hands-on IC