Software Engineering Group Manager - Site Reliability Center
Core
Lead mission-critical production operations and site reliability engineering (SRE) for enterprise platforms powering real customer experiences, ensuring infrastructure, platforms, and applications adhere to reliability standards.
Role type
Senior Group Manager, Site Reliability Engineering
Builds
Stable, reliable, and secure enterprise platforms for PNC customers
Domain
Banking/Finance + Site Reliability Engineering/Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Incident management, Root cause analysis, Change governance, Team leadership, On-call leadership, Observability strategy, Automation initiatives, Cross-functional collaboration, Production support, System design partnership
Preferred skills
Linux/Windows administration, Database management (Oracle, SQL, MongoDB, Cassandra), Middleware knowledge (Elasticsearch, Redis, MQ, Kafka), OCP experience
Technologies
Dynatrace, BigPanda, Logscale, Linux, Windows, Oracle, SQL, MongoDB, Cassandra, Elasticsearch, Redis, MQ, Kafka, OCP
Responsibilities
Lead major incident response for high-impact events; Provide technical leadership in production support and troubleshooting; Drive root cause analysis and permanent systemic solutions; Define and evolve reliability strategy across availability and resiliency; Modernize operations with best-in-class monitoring and alerting; Ensure safe and reliable change through governance and release management; Lead a Global 24x7 Operation managing distributed teams.
Seniority
Senior, hands-on IC with management scope