Senior Site Reliability Engineer
Core
Build, transition, and operate a multi-cloud production platform supporting GPO products.
Role type
Senior Site Reliability Engineer (multi-cloud)
Builds
Multi-cloud production platform (Azure, GCP, AWS) and shared platform services
Domain
Cloud infrastructure and Kubernetes
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Linux fundamentals, networking, Kubernetes, infrastructure as code, GitOps, CI/CD, observability, incident response, automation scripting
Preferred skills
Managed Kubernetes platforms (AKS, GKE, EKS), shared platform services, database technologies (MongoDB, PostgreSQL, MySQL), brownfield platform migrations
Technologies
Azure, GCP, AWS, Kubernetes, Terraform/OpenTofu, Terragrunt, Ansible, Python, Go, shell scripting, AKS, GKE, EKS
Responsibilities
Partner with architects to ensure deployable, secure, and recoverable platform designs; Assess and document multi-cloud and Kubernetes estate; Build and maintain cloud infrastructure using IaC; Operate and improve AKS, GKE, and EKS environments; Develop and support GitOps and CI/CD workflows; Troubleshoot production issues across cloud, network, and platform layers; Reduce operational toil through automation and improve SLOs, alerts, and runbooks
Seniority
Senior, hands-on IC
