Principal System Engineering - SRE
Core
Analyze production incidents end-to-end across applications, infrastructure, and cloud environments to identify root causes and prevent recurrence using observability data and AI-assisted analysis.
Role type
Principal System Engineering - SRE
Builds
Proactive reliability and incident prevention strategies for enterprise communications systems
Domain
Telecommunications / Enterprise IT Operations
Deliverable
production ML models | dashboards & analysis
Required skills
Deep RCA for production incidents, End-to-end system architecture knowledge, Observability tools expertise, Pattern identification, Postmortem writing, Operational data analysis, Distributed systems understanding, ITSM/Release Management, Python, SQL, Power BI
Preferred skills
QA/Test Engineering background, Gen AI use case building, Jira Align/JSM, SAFe/Agile/DevOps certifications
Technologies
T-APM, T-Trace, CatchPoint, Grafana, ServiceNow, Jira Cloud, Git, Power BI, Tableau, Python, SQL
Responsibilities
Analyze incidents to identify root causes and systemic weaknesses, Write high-quality postmortems, Partner with engineering teams to drive corrective actions, Implement permanent fixes and preventive improvements, Use AI/advanced analytics for incident analysis
Seniority
Principal, hands-on IC with strategy & mentorship