CaaS Private Site Reliability Engineer - Assistant Vice President
Core
Operate and improve an on-prem, multi-tenant Kubernetes platform running on bare metal to support critical, low-latency, and regulated workloads.
Role type
Senior IC Site Reliability Engineer (Kubernetes/Infrastructure)
Builds
On-prem Kubernetes clusters, observability stacks, and automation workflows for critical business services
Domain
Banking / Cloud Infrastructure / Kubernetes
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Kubernetes administration, Linux system administration, Python scripting, Ansible, Bash, observability stack management, incident management, root cause analysis, computer networking, virtualization, containerization, distributed systems
Preferred skills
Service mesh (Istio/Envoy), stateful services (PostgreSQL/Kafka/MongoDB), chaos testing, low-latency environment experience, Golang code reading
Technologies
Kubernetes, Prometheus, Grafana, Splunk, OpenTelemetry, Ansible, Python, Bash, Istio, Envoy, PostgreSQL, Kafka, MongoDB
Responsibilities
Define and improve SLI/SLOs and error budgets; Build and maintain observability across metrics, logs, and alerts; Lead incident response and postmortems; Automate operational tasks and remediation workflows; Improve reliability and upgrade safety for Kubernetes clusters and dependencies; Partner with teams on capacity planning and release readiness
Seniority
Senior, hands-on IC
