Site Reliability Engineer (AIOps) (f/m/d) @ A1 Competence Delivery Center
Core
Build and optimize a multi-cloud monitoring ecosystem using AIOps to ensure end-to-end visibility and proactive issue identification during infrastructure transformation.
Role type
Site Reliability Engineer (AIOps)
Builds
Multi-cloud observability pipelines and automated self-healing solutions for Azure and Exoscale environments.
Domain
Telecommunications / Cloud Infrastructure / AIOps
Deliverable
production ML models | infrastructure
Required skills
DevOps/SRE experience, Kubernetes, observability stack (Prometheus, Grafana, OpenTelemetry, ELK/PLG), AIOps platforms (Datadog, Dynatrace, New Relic), Python/R, cloud migration expertise
Preferred skills
Experience with GitHub Actions CI/CD integration, custom monitoring solution development
Technologies
Azure, Exoscale, Kubernetes, Prometheus, Grafana, OpenTelemetry, ELK, PLG, Datadog, Dynatrace, New Relic, GitHub Actions, Python, R
Responsibilities
Design and optimize multi-cloud observability pipelines for metrics, logs, and traces; Implement AIOps solutions for anomaly detection and event correlation; Ensure seamless monitoring during cloud migration; Integrate observability into CI/CD pipelines; Develop automation and self-healing solutions with DevOps engineers
Seniority
Mid-to-Senior, hands-on IC