Senior Systems Operations Engineer
Core
End-to-end operational health and reliability of enterprise monitoring, scheduling, and observability platforms.
Role type
Senior Systems Operations Engineer (IC)
Builds
Production support, incident response, platform health, upgrades, and operational readiness for critical enterprise platforms.
Domain
Enterprise IT Operations, Observability, Scheduling
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
ITIL execution, SRE principles, reliability engineering, automation, root cause analysis, change management, incident triage, observability practices, runbook creation, on-call response
Preferred skills
Enterprise application platform support, large-scale environment experience, independent operation, cross-timezone collaboration
Technologies
Autosys, Grafana, Splunk, Cribl, ThousandEyes
Responsibilities
Provide hands-on production support for monitoring, scheduling, and observability platforms; Ensure platforms meet availability, reliability, and performance objectives; Drive incident management activities including triage, escalation, and remediation; Contribute to problem management by identifying recurring issues and driving root-cause fixes; Execute change management activities such as patching, upgrades, and configuration changes; Proactively monitor platform health and identify risks to stability, capacity, or performance; Create, maintain, and enhance runbooks and playbooks; Participate in on-call rotations and provide timely response during production incidents
Seniority
Senior, hands-on IC