Staff Infrastructure Engineer
Core
Staff Infrastructure Engineer leveraging agentic AI to identify, assess, and resolve critical service issues and major incidents within the Technology Command Center.
Role type
Staff Infrastructure Engineer (SRE/Reliability)
Builds
Enterprise resilience and proactive AI-driven operations
Domain
Insurance technology, Cloud infrastructure, Observability
Deliverable
production ML models | infrastructure
Required skills
AI-driven data analysis, Prompt engineering, Alert triage, Event correlation, Service impact analysis, Automation workflow enablement, Monitoring tool management, Incident escalation, Runbook execution, Cross-functional collaboration
Preferred skills
None stated
Technologies
Splunk, Dynatrace, ITSI, Moogsoft, ThousandEyes
Responsibilities
Manage technical and executive communication for major incidents; Perform alert triage, correlation, and initial impact assessment; Support major incident detection and escalation by validating symptoms and engaging resolver teams; Use standard operating procedures to investigate, prioritize, and escalate events; Maintain situational awareness during active events; Identify opportunities to improve monitoring effectiveness, event quality, automation, and operational readiness
Seniority
Staff, hands-on IC with strategic partnership