AI Ops Engineer (m/w/d)
Core
Design and implement AI-driven automation, observability, and ML-based system analysis to build highly available, robust, and autonomous IT operations platforms.
Role type
Senior IC AI Ops Engineer
Builds
Intelligent, scalable IT operations solutions including automated remediation, self-healing mechanisms, and predictive analytics systems.
Domain
IT Operations / DevOps / Machine Learning
Deliverable
production ML models | infrastructure
Required skills
Python for automation and data analysis, Observability stacks (Prometheus, Grafana, ELK/Elastic, OpenTelemetry), Big-Data/Streaming technologies (Kafka, Spark), Cloud platforms (AWS/Azure/GCP) for monitoring and telemetry, Container orchestration (Podman, Kubernetes), Infrastructure automation (Ansible, Terraform, Argo, GitOps), Root Cause Analysis (RCA), Event correlation and noise reduction, Anomaly detection, Predictive analytics, Automated remediation, Self-healing mechanisms, Scripting, Incident response and troubleshooting.
Preferred skills
Technical education in Informatics, Data Science, or equivalent practical experience.
Responsibilities
Develop AI-supported automation solutions to make IT operations more efficient, stable, and proactive; Implement event correlation and noise reduction mechanisms to detect important signals quickly; Develop anomaly detection and predictive analytics systems for early prediction of critical failures; Conduct root cause analyses using modern ML and data analysis methods; Implement automated remediation and self-healing mechanisms to automatically fix or mitigate incidents; Build a continuous observability architecture (Logs, Metrics, Traces) to ensure full system transparency; Automate recurring processes using scripting and infrastructure automation; Support incident response and troubleshooting; Collaborate with AI engineers, platform teams, and infrastructure experts to develop intelligent, scalable operations solutions.
Seniority
Senior, hands-on IC