Associate Director Engineering - SRE
Core
Lead Site Reliability Engineering (SRE) initiatives to ensure the health, availability, and continuous improvement of critical production environments for HSBC's Global Service Center.
Role type
Senior IC SRE with management responsibilities (Associate Director)
Builds
Scalable, highly available, and secure cloud infrastructure and services
Domain
Financial Services / Cloud Infrastructure / Site Reliability Engineering
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Cloud platform expertise (AWS/GCP/Azure), Container orchestration (Kubernetes, Docker), SRE methodologies (SLIs/SLOs, error budgets), Observability (Prometheus, Grafana, ELK, Datadog, Splunk), CI/CD automation, Infrastructure as Code (Terraform, Ansible, Helm), Incident management, Mentoring
Preferred skills
Microservices and distributed systems architecture, AIOps, SRE and cloud certifications (GCP Professional SRE, AWS DevOps Engineer, CKA/CKAD), 24x7 high-availability environment experience
Technologies
AWS, GCP, Azure, Kubernetes, Docker, Prometheus, Grafana, ELK, Datadog, Splunk, Terraform, Ansible, Helm
Responsibilities
Lead complex troubleshooting and root cause analysis for production incidents; Design and architect scalable cloud infrastructure; Champion SRE practices and automation; Develop monitoring and alerting systems; Mentor junior engineers; Collaborate with development and operations teams; Lead on-call rotations and postmortems; Drive large-scale system improvements
Seniority
Senior, hands-on IC with management/mentorship duties