Observability Platform Engineer
Core
Lead and evolve the observability strategy for cloud and on-premises environments, serving as the primary owner of the Datadog platform to ensure 24/7 service health validation across business-critical systems.
Role type
Senior IC Observability Platform Engineer
Builds
Scalable monitoring solutions, dashboards, SLOs, and alerting frameworks using Datadog across Windows and Linux/Unix environments.
Domain
Financial services / Cloud Infrastructure / Observability
Deliverable
production ML models | product features | dashboards & analysis
Required skills
Datadog platform expertise (APM, Logs, Traces, Synthetics, RUM, SLOs), Python scripting, Terraform, distributed tracing, SLO/error budget frameworks, ITSM integrations (ServiceNow), cloud platform integration (Azure/AWS), agent deployment, OS-level performance analysis.
Preferred skills
Legacy monitoring migration experience, .NET instrumentation, network monitoring concepts, CI/CD pipeline integration.
Technologies
Datadog, Terraform, Python, PowerShell, Bash, OpenTelemetry, ServiceNow, Azure, AWS, OpenShift, Windows Server, Linux, Solaris.
Responsibilities
Architect and maintain scalable observability solutions; lead migration from OpenView to Datadog; automate configuration management via APIs and IaC; define and operationalize SLOs; integrate with ServiceNow for incident routing; champion OpenTelemetry adoption; onboard new applications; optimize platform cost and scaling.
Seniority
Senior, hands-on IC