Observability Platform Engineer (Datadog)
Core
Lead and evolve the observability strategy for cloud and on-premises environments, serving as the primary owner of the Datadog platform to ensure 24/7 service health validation across business-critical systems.
Role type
Senior individual contributor platform engineer (Datadog)
Builds
End-to-end observability solutions including logs, metrics, traces, SLOs, synthetic monitoring, and Real User Monitoring (RUM) for reliability and incident response.
Domain
Cloud infrastructure and on-premises systems monitoring
Required skills
Datadog platform expertise, distributed tracing, SLO/error budget frameworks, Terraform, Python, PowerShell, Bash, Windows Server, Unix/Linux/Solaris, ITSM integrations
Preferred skills
Legacy monitoring migration, .NET instrumentation, financial services experience, CI/CD integration, network monitoring concepts
Technologies
Datadog, Terraform, ServiceNow, OpenTelemetry, Azure, AWS
Responsibilities
Architect scalable observability solutions across cloud and on-prem environments; lead migration from OpenView to Datadog; automate configuration management via APIs and scripting; implement APM, tracing, log management, and NPM; define and operationalize SLOs; integrate with ServiceNow for incident routing; champion OpenTelemetry adoption; onboard new applications and provide engineering enablement.
Seniority
Senior, hands-on IC