Director Site Reliability Engineering
Core
Lead the transformation of reliability, performance, and availability across platforms by modernizing operational practices through automation, cloud-native design, and API-driven integration.
Role type
Director, Site Reliability Engineering
Builds
AWS cloud architecture and MuleSoft integration ecosystem
Domain
Financial Services / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
AWS services (ECS, EKS, S3, RDS, VPC), MuleSoft, API/Integrations (Apigee), Terraform, GitLab, Okta, OAuth2, Dynatrace, Python, containerization, SLA/SLO/SLI definition, cloud migration
Preferred skills
Agile methodology (SAFe, Scrum), Jira, Confluence, unit testing frameworks, RESTful APIs, micro-services, WCF, TSQL, SQL, application security tools (Owasp, Veracode, AppScan)
Technologies
AWS, MuleSoft, Apigee, Terraform, GitLab, Okta, Dynatrace, Python, Jira, Confluence
Responsibilities
Implement and maintain monitoring, logging, and tracing tools; write software and scripts to automate deployment, monitoring, and system management; respond to incidents and perform root cause analysis; design and build reliable, scalable systems and define SLOs/SLIs; collaborate with developers to ensure application reliability; create and maintain runbooks and system diagrams; partner with peers to advance DevOps maturity and set technical direction; provide leadership and technical expertise to Agile teams for sprint planning and release validation; ensure quality, performance, and security of systems meet customer SLAs and audit expectations.
Seniority
Director, strategic leadership with hands-on technical execution