AI Ops Engineer
Core
Design and maintain intelligent monitoring and incident response workflows; build data pipelines for operational analytics; develop predictive models for anomaly detection and outage forecasting.
Role type
Senior IC AI Ops Engineer
Builds
Intelligent monitoring workflows, data pipelines, and predictive models for system reliability
Domain
Cloud and on-premises IT operations, observability, and AI-driven automation
Deliverable
production ML models | product features
Required skills
Python, observability platforms (Splunk, Datadog, Prometheus, Grafana), cloud platform operations (Azure, AWS, Google Cloud), IT service management, workflow orchestration
Preferred skills
Machine learning model lifecycle management and deployment
Technologies
Python, Splunk, Datadog, Prometheus, Grafana, Azure, AWS, Google Cloud, ServiceNow, PagerDuty, Rundeck
Responsibilities
Design and maintain intelligent monitoring and incident response workflows across cloud and on-premises systems; Build and optimize data pipelines to collect, normalize, and enrich logs, metrics, and events; Develop and deploy predictive models to detect anomalies, forecast outages, and reduce mean time to resolution; Automate remediation runbooks and integrate alerting with service management platforms; Collaborate with DevOps, Site Reliability, and security teams to improve platform performance and operational governance
Seniority
Mid-to-Senior, hands-on IC
