Manager, AI Ops Site Reliability Engineer
Core
Design, build, and evaluate AI agents and workflows to automate IT operations, incident investigation, and knowledge retrieval for Pfizer's global infrastructure.
Role type
Manager, AI Ops Site Reliability Engineer
Builds
AI-driven operational agents, evaluation harnesses, and automated incident response flows
Domain
IT Operations / Site Reliability Engineering / Generative AI
Deliverable
production ML models | product features
Required skills
LLM and agent-based application development, prompt and context engineering, structured evaluation of AI systems (ground-truth sets, metrics, regression testing), agent frameworks (Amazon Bedrock AgentCore, LangGraph), RAG and knowledge-base design, Python, IT operations literacy
Preferred skills
LLM-as-judge evaluation, observability tools (ServiceNow, Dynatrace)
Technologies
Amazon Bedrock AgentCore, LangGraph, Python
Responsibilities
Design agent workflows for intake, investigation, and decision logic; build and own the evaluation harness for accuracy and hallucination detection; define reasoning guardrails and confidence thresholds; convert tacit operational knowledge into machine-usable context; pilot and measure toil/MTTR reduction in operational workflows; continuously tune agent quality against production results
Seniority
Manager, hands-on IC with team leadership
