CaaS Private Site Reliability Lead Engineer - Vice President
Core
Lead reliability, resilience, and operational excellence for the CaaS Private platform in the US, focusing on Kubernetes, observability, automation, and incident management.
Role type
Senior IC Site Reliability Engineer (Kubernetes/Cloud-Native)
Builds
Production-ready CaaS Private platform services for Deutsche Bank's US business areas
Domain
Financial Services / Cloud-Native Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Kubernetes (Bare Metal), Linux, distributed systems reliability, observability, monitoring, alerting, incident response, root cause analysis, automation, self-healing workflows, SLO definition, capacity planning, disaster readiness, AI tool integration
Preferred skills
Infrastructure-as-code, platform architecture influence, strategic communication, mentoring, operational judgment
Technologies
Kubernetes, Linux, AI tools
Responsibilities
Lead reliability strategy including SLO frameworks and incident management maturity; Drive resilience improvements in observability and automation; Own complex production troubleshooting and preventive fixes; Define service indicators, alert thresholds, and production readiness criteria; Develop self-healing workflows to reduce manual intervention; Mentor engineers and foster a culture of blameless learning
Seniority
Senior, hands-on IC with leadership responsibilities