AI Observability Engineer
Core
Design, implement, and support enterprise observability platforms while driving automation, AI, and Agentic AI initiatives to improve service reliability and reduce operational overhead.
Role type
Senior IC observability engineer (AI/AIOps)
Builds
Enterprise observability platforms, automated remediation workflows, AI-driven anomaly detection, and self-healing solutions
Domain
Unified commerce, cloud infrastructure, SRE, AIOps
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Observability platform design, scripting (Python/PowerShell/Bash), cloud platforms (Azure/AWS/GCP), CI/CD, Infrastructure as Code, incident management, root cause analysis, AI-assisted engineering practices
Preferred skills
Microsoft Copilot, Azure OpenAI, Copilot Studio, AIOps platforms, OpenTelemetry, Kubernetes, LangChain, Semantic Kernel, Agentic AI frameworks, RAG architectures, vector databases, prompt engineering
Technologies
Splunk, AppDynamics, Splunk Observability Cloud, Microsoft Azure, AWS, Google Cloud, Terraform, Bicep, ARM, OpenTelemetry, Kubernetes, LangChain, Semantic Kernel
Responsibilities
Design and deploy monitoring, logging, and tracing platforms; develop dashboards and KPIs for operational insights; automate alert management and incident response workflows; implement AI-powered anomaly detection and predictive analytics; build Agentic AI solutions for autonomous alert analysis and remediation; collaborate with engineering teams to improve application performance and scalability; participate in on-call support and major incident management
Seniority
Senior, hands-on IC