Senior Observability Engineer
Core
Design and support an agentic infrastructure observability platform that investigates incidents across live topology, telemetry, logs, and events for large datacenters and telecom environments.
Role type
Senior IC observability engineer (LLM orchestration & time-series ML)
Builds
An agentic infrastructure observability platform for large datacenters and telecom environments
Domain
Telecom infrastructure, datacenter operations, observability, and AI-driven incident response
Deliverable
production ML models | product features
Required skills
Python (async, FastAPI, Pydantic), LLM orchestration (DSPy, LangGraph, LangChain), RAG system design, PromQL/Prometheus, time-series ML (anomaly detection, forecasting), graph databases (Neo4j, Cypher), Docker/Kubernetes deployment
Preferred skills
None stated
Technologies
asyncio, httpx, FastAPI, Pydantic v2, DSPy, LangGraph, LangChain, PromQL, Prometheus, Grafana, Neo4j, Docker, Kubernetes, SONiC, Cumulus, iDRAC, MCP, OpenAPI, Swagger
Responsibilities
Design and support REST and async API layers with versioned endpoints and authentication patterns; build and refine DSPy agent programs for investigation flows from alert intake to remediation proposals; design and implement MCP tool servers exposing Neo4j infrastructure graphs; implement propose-approve-apply governance flows with audit ledgers; write PromQL for recording and alerting rules; apply machine learning to network and server telemetry for anomaly detection and forecasting; build RAG layers over operational logs and vendor documentation; create telemetry ingestion pipelines from SONiC, Cumulus, and iDRAC into Prometheus, Loki, and Neo4j; build evaluation harnesses and behavioral regression tests for agents; containerize and deploy services with Docker and Kubernetes; mentor junior engineers on testing, observability, and AI governance
Seniority
Senior, hands-on IC