Sr. Platform Operations Engineer
Core
Identify monitoring and performance optimization opportunities, apply AI/ML to anomaly detection and alert correlation, and conduct incident root cause analysis to improve platform stability.
Role type
Senior IC Platform Operations Engineer (AIOps)
Builds
AI-assisted tooling, automation workflows, and proactive incident response strategies
Domain
Cloud infrastructure, observability, and AI-driven operations
Deliverable
production ML models
Required skills
Anomaly detection, alert correlation, incident root cause analysis, APM platform administration, log management, scripting (PowerShell/Python), Configuration-As-Code, containerized workload monitoring
Preferred skills
Payment processing experience, PCI policy knowledge, AI agent/LLM tooling experience, mentoring
Technologies
Dynatrace, Datadog, New Relic, Splunk, Azure Logic Apps, Azure Functions, Azure AI Foundry, PowerShell, Python
Responsibilities
Implement monitoring and performance optimizations, apply AI/ML for operational issue resolution, document SOPs and alert-response procedures, serve as escalation point for complex stability issues, build AI-assisted automation to reduce manual operations
Seniority
Senior, hands-on IC