Principal Support Engineer – Customer Reliability & Escalations Engineering
Core
Senior technical authority leading critical customer escalations, driving root cause analysis across distributed systems, and influencing platform reliability for strategic enterprise clients.
Role type
Principal Support Engineer (Customer Reliability & Escalations)
Builds
High-availability platform services and integrations for global enterprise technology distribution
Domain
Enterprise technology distribution, SaaS, cloud infrastructure, and distributed systems
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Incident leadership, root cause analysis, distributed systems troubleshooting, cloud platform expertise, cross-functional collaboration, technical advisory, mentorship
Preferred skills
SRE practices, DevOps workflows, CI/CD pipeline management, scripting (Python, Bash, PowerShell), fintech domain knowledge
Technologies
AWS, Azure, GCP, Kubernetes, containers, Linux, Datadog, Dynatrace, Splunk, Grafana, New Relic, AppDynamics, Elastic, REST APIs, microservices
Responsibilities
Lead resolution of Sev1 and Sev2 incidents impacting strategic customers, drive end-to-end troubleshooting across applications and cloud infrastructure, perform deep root cause analysis using logs and metrics, partner with Engineering to implement permanent corrective actions, identify recurring issues to improve reliability and automation, serve as trusted technical advisor during executive-level escalations, mentor engineers and promote best practices in incident management
Seniority
Principal, strategic leadership & hands-on technical authority