Site Reliability Engineer - Application Support (Director)
Core
Director-level SRE leading production stability, outage management, and optimization for Wealth Management application platforms.
Role type
Director, Site Reliability Engineer (Application Support)
Builds
Wealth Management Investment Management application platforms and production environment stability
Domain
Financial Services / Wealth Management Technology
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Outage management, incident/problem/change management, cross-team coordination, technical architecture documentation, high-availability system administration, database engineering, scripting (Python, Java, Perl, PowerShell, Unix), cloud services (AWS EC2, ECS, S3, Fargate, Aurora, Lambda), observability tools (Grafana, Prometheus, Splunk, Kibana), Agile/DevOps practices
Preferred skills
Financial services industry experience, web analytics tools (Adobe Experience Cloud), mentoring/coaching team members
Technologies
AWS (EC2, ECS, S3, Fargate, Aurora, Lambda), Python, Java, Perl, PowerShell, Unix, JavaScript, React, GraphQL, Django, Celery, PostgreSQL, Golang, ElasticSearch, RabbitMQ, Kafka
Responsibilities
Proactively detect, troubleshoot, and resolve production application issues; maintain clear communications during outages; develop and revise policies/procedures for production standards; gatekeep change implementation; service data/access requests for production systems; collaborate with development teams on new system standards; maintain a body of knowledge for team self-reliance; mentor and coach team members
Seniority
Director, strategic leadership with hands-on technical expertise
