CareerPlanSign in

Site Reliability Engineer, AI Observability

New York💼 Full-time🗓 2026-10-06 → 2026-10-07

Core

Build and operate a shared observability platform for LLM-based applications, establishing monitoring, alerting, and service levels from scratch.

Role type

Senior Site Reliability Engineer (AI Observability)

Builds

Self-hosted observability platform on AWS (Kubernetes, ClickHouse, PostgreSQL, Redis) for a global life sciences client (via careerplan.io/jobs/8870682002-site-reliability-engineer-ai-observability-at-appnovation)

Domain

Cloud Infrastructure, AI Observability, Life Sciences

Deliverable

production ML models

Required skills

Kubernetes (EKS), Helm, Argo CD, ClickHouse, PostgreSQL, Redis, Infrastructure as Code, Runbook/SOP creation, Incident management, Automation

Preferred skills

OpenTelemetry, Grafana, OIDC/SSO, GitHub Actions, LLM observability tools (Langfuse, LangSmith, Arize Phoenix), Life Sciences industry experience

Technologies

AWS, Kubernetes, Helm, Argo CD, ClickHouse, PostgreSQL, Redis, Langfuse, OpenTelemetry, Grafana, GitHub Actions

Responsibilities

Diagnose and resolve failures across the stack including ClickHouse, PostgreSQL, Redis, and Kubernetes; Own ClickHouse operator model including replication, topology, and backups; Build monitoring, alerting, and service levels from scratch; Maintain Infrastructure as Code (manifests, Helm, Argo CD); Write and maintain runbooks and SOPs for incident resolution; Manage onboarding and support for internal teams; Automate recurring operational tasks; Partner with platform engineers on safe upgrades and migrations

Seniority

Senior, hands-on IC