AI Observability Engineer
Core
Own the AI observability backbone for an Azure AI cloud platform, ensuring LLM and agent tracing, platform metrics, and infrastructure health are visible, measurable, and reliable.
Role type
Senior IC AI Observability Engineer
Builds
LLM and agent monitoring tooling, Grafana dashboards, Prometheus metrics, and Azure observability stack
Domain
Cloud Infrastructure / AI Platform Engineering
Deliverable
production ML models | infrastructure
Required skills
Azure Monitor, Application Insights, Log Analytics, Managed Grafana, Langfuse, Prometheus, Terraform, CI/CD, Python, ML workload instrumentation, observability fundamentals (logs, metrics, traces, alerting, SLIs, SLOs)
Preferred skills
PromQL, Kusto Query Language, OpenTelemetry, GenAI semantic conventions, LLM evaluation frameworks, AI cost dashboards, AKS and Kubernetes observability
Technologies
Azure, Langfuse, Grafana, Prometheus, Terraform, Python, Kubernetes
Responsibilities
Stand up and operate LLM and agent monitoring with Langfuse; Capture traces, latency, token usage, cost, quality scores, prompt and model-version analytics, and safety signals; Build lightweight internal tooling and exporters in Python; Design and maintain Grafana dashboards, Prometheus metrics, and the Azure observability stack; Instrument platform and AI workloads for health, usage, cost, and SLA reporting; Feed telemetry and operational insights into the Platform Engineering backlog; Own Terraform IaC and CI/CD for observability tooling; Support incident investigation and root-cause analysis
Seniority
Senior, hands-on IC