Lead Platform Reliability Engineer
Core
Lead Platform Reliability Engineer specializing in Network, Middleware, Database, or Storage to ensure enterprise platform stability, resiliency, and operational excellence.
Role type
Senior IC Platform Reliability Engineer (Infrastructure/SRE)
Builds
Critical enterprise platforms (Network, Middleware, Database, Storage) for Wells Fargo
Domain
Financial Services / Enterprise Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Systems Engineering, Enterprise-scale production environment support, Network Engineering, Middleware Engineering, Database Engineering, Storage Engineering, SRE principles (SLIs/SLOs/Error Budgets), Capacity planning, Performance analysis, Observability (metrics/logging/tracing), Automation/Scripting, Infrastructure as Code
Preferred skills
Troubleshooting complex multi-domain issues, Disaster recovery, Self-healing capabilities, CI/CD pipelines, Mentoring engineers
Technologies
Python, Bash, PowerShell, Grafana, Splunk, Prometheus, AppDynamics, Cribl, ThousandEyes, Dynatrace, Ansible, Terraform, WebSphere, Tomcat, JBoss, Kafka, MQ, Oracle, SQL Server, PostgreSQL, MongoDB, SAN/NAS
Responsibilities
Lead investigation and resolution of complex production incidents, Apply SRE principles to improve platform health, Lead capacity analysis and forecasting, Perform deep performance analysis across infrastructure layers, Identify and remediate configuration drift and operational debt, Design and implement automation solutions to reduce toil, Define and enhance enterprise observability standards, Lead blameless post-incident reviews, Mentor engineers on reliability engineering best practices
Seniority
Senior, hands-on IC