Site Reliability Engineer, AI Observability
Core
Build and operate a shared observability platform for LLM-based applications, establishing monitoring, alerting, and service levels from scratch.
Role type
Senior Site Reliability Engineer (AI Observability)
Builds
Self-hosted observability platform on AWS (Kubernetes, ClickHouse, PostgreSQL, Redis) for a global life sciences client (via careerplan.io/jobs/8870682002-site-reliability-engineer-ai-observability-at-appnovation)
Domain
Cloud Infrastructure, AI Observability, Life Sciences
Deliverable
production ML models
Required skills
Kubernetes (EKS), Helm, Argo CD, ClickHouse, PostgreSQL, Redis, Infrastructure as Code, Runbook/SOP creation, Incident management, Automation
Preferred skills
OpenTelemetry, Grafana, OIDC/SSO, GitHub Actions, LLM observability tools (Langfuse, LangSmith, Arize Phoenix), Life Sciences industry experience
Technologies
AWS, Kubernetes, Helm, Argo CD, ClickHouse, PostgreSQL, Redis, Langfuse, OpenTelemetry, Grafana, GitHub Actions
Responsibilities
Diagnose and resolve failures across the stack including ClickHouse, PostgreSQL, Redis, and Kubernetes; Own ClickHouse operator model including replication, topology, and backups; Build monitoring, alerting, and service levels from scratch; Maintain Infrastructure as Code (manifests, Helm, Argo CD); Write and maintain runbooks and SOPs for incident resolution; Manage onboarding and support for internal teams; Automate recurring operational tasks; Partner with platform engineers on safe upgrades and migrations
Seniority
Senior, hands-on IC