Senior Manager, AI Reliability Engineering -Kroger Technology & Digital (P2498)
Core
Leads the engineering discipline to make enterprise AI operationally trustworthy at scale, defining production-grade standards and building reliability tooling for models, agents, and copilots.
Role type
Senior Manager, AI Reliability Engineering
Builds
Enterprise AI reliability standards, automation tooling, observability systems, and production-readiness frameworks for models and agents.
Domain
Retail technology, Enterprise AI, Cloud Infrastructure
Deliverable
production ML models
Required skills
Distributed systems, Kubernetes, Cloud infrastructure (GCP/Azure), SRE practices, ML/AI/LLM production experience, Incident management, Technical leadership
Preferred skills
LLMOps, Model observability, Evaluation tooling (LangSmith, MLflow, Arize, Fiddler)
Technologies
Kubernetes, GCP, Azure, LangSmith, MLflow, Arize, Fiddler
Responsibilities
Define production-grade AI standards and quality bars; Stand up the AI Reliability Engineering function; Partner with teams to embed resilience and observability into AI systems; Own live observability and production-quality signals for models and agents; Drive efficiency by optimizing inference cost and token usage; Lead incident response and establish agentic incident prevention practices; Hire and grow a multidisciplinary team.
Seniority
Senior, hands-on IC with leadership responsibilities
