Senior DevOps / Observability Engineer – AI Runtime & Platform Monitoring
Core
Build and maintain observability platforms for modern cloud-native and AI-assisted platforms, focusing on logging, tracing, telemetry, and runtime monitoring of AI workloads and agentic services.
Role type
Senior DevOps / Observability Engineer (AI Runtime & Platform Monitoring)
Builds
Observability platforms, logging/metrics/tracing pipelines, and monitoring solutions for distributed systems and AI services.
Domain
Cloud-native infrastructure, AI platforms, and distributed systems operations.
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Kubernetes, Prometheus, Grafana, OpenTelemetry, ELK/OpenSearch, CI/CD, Linux, automation, incident handling, distributed tracing, metrics collection, log analysis.
Preferred skills
AI observability, telemetry pipelines for AI systems, tracing agentic workflows, runtime monitoring of AI platforms.
Technologies
Kubernetes, Prometheus, Grafana, OpenTelemetry, ELK, OpenSearch, Linux.
Responsibilities
Build and maintain observability platforms; work with logging, metrics, and distributed tracing; monitor cloud-native and AI-related workloads; support stable operations and incident response; optimize performance and reliability in production.
Seniority
Senior, hands-on IC