Site Reliability Engineer, IaaS
Core
Building the foundations of a unified cloud and Kubernetes platform to replace a legacy bare-metal fleet, enabling safe workload migration and operations.
Role type
Senior IC Site Reliability Engineer (IaaS)
Builds
Cloud baseline capabilities, Kubernetes infrastructure, automation, and self-service platform modules
Domain
Cloud Infrastructure / Kubernetes / Site Reliability Engineering
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
AWS or GCP production experience, Kubernetes operations, Infrastructure as Code (Terraform), Python or Go programming, Linux fundamentals, networking fundamentals
Preferred skills
Multi-cloud experience, GitOps or policy-as-code tooling (Argo CD, Helm, OPA, Kyverno), cloud migration experience, platform engineering experience
Technologies
AWS, GCP, Kubernetes, Terraform, Python, Go, GitOps, Argo CD, Helm, OPA, Kyverno
Responsibilities
Build and improve Cloud Baseline capabilities including identity, access, networking, security, and resource inventory; Develop and maintain infrastructure as code and automation for cloud and Kubernetes environments; Contribute to reliable, repeatable cloud and cluster lifecycle operations; Build self-service capabilities, reusable modules, and documentation; Reduce manual work and configuration drift through automation, testing, and GitOps practices; Improve observability, monitoring, alerting, and capacity management; Investigate production issues and participate in on-call rotation
Seniority
Senior, hands-on IC