Senior AI Platform Operations Engineer
Core
Ensure reliability, operability, and controlled enablement of the organization's AI platform, including production readiness, monitoring, incident coordination, and governance enforcement.
Role type
Senior AI Platform Operations Engineer (SRE/DevOps)
Builds
Enterprise AI platform services, Copilot integrations, and AI-driven operational solutions
Domain
Enterprise IT Operations / Cloud / AI Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Cloud platform administration, observability, automation, incident management, governance control execution, ITIL/ITSM processes, root-cause analysis, infrastructure-as-code
Preferred skills
AI/ML operational concepts, human-in-the-loop practices, enterprise network security, cost optimization
Technologies
Azure Automation, Azure AI, Copilot Studio, AKS, Azure Monitor, Application Insights, Grafana, Power Platform, Bicep, Terraform, Azure Policy, Service Bus, Event Grid, Apache Kafka, Elastic, Azure AI Search, Cosmos DB
Responsibilities
Administer and operate the AI platform to ensure availability and resilience; Monitor platform health and coordinate incident response; Enable approved AI use cases into production; Implement observability capabilities and AI Ops use cases; Execute governance controls for AI solutions; Maintain operational visibility of AI platform assets
Seniority
Senior, hands-on IC