CareerPlanSign in

Observability & AgentOps Engineer

💼 Full-time🗓 2026-10-04

Core

Design and run the systems that keep AI agents reliable in production, focusing on monitoring, deployment, tracing, and recovery.

Role type

Senior IC Observability & AgentOps Engineer

Builds

Production-ready agent infrastructure including tracing, logging, deployment pipelines, and recovery systems.

Domain

Financial Services & Insurance / AI Agent Operations

Deliverable

production ML models

Required skills

distributed tracing, metrics, structured logging, SLIs/SLOs definition, alerting design, CI/CD, Docker, AWS Step Functions, Python 3.11+

Preferred skills

LLM/agent workload design, infrastructure-as-code, high-volume log/trace storage design

Technologies

Langfuse, OpenTelemetry, ELK, AWS Step Functions, Docker, Python

Responsibilities

Design end-to-end observability architecture (tracing, metrics, logging); Define reliability targets (SLIs/SLOs) and alerting; Design deployment topology and orchestration; Design failure recovery mechanisms; Design log and trace data model and storage; Set up day-to-day monitoring and support handover.

Sourced via wellfound · Listed on CareerPlan, which tracks 940,000+ jobs from 20+ sources.