CareerPlanSign in

AI Observability Engineer

💼 Full-time🗓 2026-08-05 → 2026-09-25

Core

Own the AI observability backbone for an Azure AI cloud platform, ensuring LLM and agent tracing, platform metrics, and infrastructure health are visible, measurable, and reliable.

Role type

Senior IC AI Observability Engineer

Builds

LLM and agent monitoring tooling, Grafana dashboards, Prometheus metrics, and Azure observability stack

Domain

Cloud Infrastructure / AI Platform Engineering

Deliverable

production ML models | infrastructure

Required skills

Azure Monitor, Application Insights, Log Analytics, Managed Grafana, Langfuse, Prometheus, Terraform, CI/CD, Python, ML workload instrumentation, observability fundamentals (logs, metrics, traces, alerting, SLIs, SLOs)

Preferred skills

PromQL, Kusto Query Language, OpenTelemetry, GenAI semantic conventions, LLM evaluation frameworks, AI cost dashboards, AKS and Kubernetes observability

Technologies

Azure, Langfuse, Grafana, Prometheus, Terraform, Python, Kubernetes

Responsibilities

Stand up and operate LLM and agent monitoring with Langfuse; Capture traces, latency, token usage, cost, quality scores, prompt and model-version analytics, and safety signals; Build lightweight internal tooling and exporters in Python; Design and maintain Grafana dashboards, Prometheus metrics, and the Azure observability stack; Instrument platform and AI workloads for health, usage, cost, and SLA reporting; Feed telemetry and operational insights into the Platform Engineering backlog; Own Terraform IaC and CI/CD for observability tooling; Support incident investigation and root-cause analysis

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.