SiteOps Engineer
Core
Managing and resolving incidents and bugs across digital platforms to ensure stability, reliability, and continuous improvement.
Role type
Site Operations Engineer (Incident Management & Platform Reliability)
Builds
Digital business platforms and operational processes
Domain
Technology Distribution / Web Solutions
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
C# (.NET Framework), ASP.NET MVC, REST APIs, Incident Management, Root Cause Analysis, Observability Tools, Log Analysis, Agile/Scrum, Stakeholder Management
Preferred skills
ITIL v4, Docker, Kubernetes, SLA Governance, Change Management
Technologies
C#, .NET, ASP.NET MVC, REST APIs, Dynatrace, New Relic, Prometheus, Datadog, Graylog, Log Analytics, AppInsights, Docker, Kubernetes
Responsibilities
Lead and manage incidents ensuring correct triage, prioritization, investigation, and resolution within SLAs; Perform root-cause analysis and collaborate on technical solutions for bugs; Identify recurring issues and platform trends to reduce incident volume; Act as change lead for bug fixes and operational improvements; Mentor and support other SiteOps team members; Maintain high-quality incident and backlog management.
Seniority
Mid-level, hands-on IC