4468 - Site Reliability Engineer III
Core
Deploy and operate hundreds of services on OpenShift in customer data centers using a GitOps-driven model, ensuring high-throughput rollouts and platform health.
Role type
Senior Site Reliability Engineer (Platform Engineering)
Builds
Production-grade on-premise OpenShift platforms with automated CI/CD pipelines and validated service deployments
Domain
Healthcare technology / Cloud Infrastructure / On-premise Kubernetes
Deliverable
production ML models | product features | infrastructure
Required skills
Kubernetes, OpenShift, GitOps (ArgoCD), Helm, CI/CD pipelines, Python/Bash/Go scripting, observability (Prometheus/Grafana), secrets management
Preferred skills
HashiCorp Vault, regulated environment experience, certificate handling
Technologies
OpenShift, ArgoCD, Helm, GitLab CI, Bitbucket Pipelines, KEDA, Vault, Prometheus, Grafana
Responsibilities
Author and maintain Helm values overlays and application manifests for OpenShift deployments; Re-point CI pipelines and service configurations from cloud to on-premise targets; Operate shared platform layer including data stores, autoscaling, and ingress gateways; Execute per-service validation including health checks and smoke tests; Build monitoring dashboards and contribute to runbooks; Troubleshoot deployment and networking issues in default-deny network models
Seniority
Senior, hands-on IC