Support Engineer, AI/ML & Platform Operations (8am-5pm) or (12md-9pm)
Core
Resolve incidents and optimize performance for Workday's enterprise SaaS platform and autonomous AI agent workflows.
Role type
Senior Support Engineer (AI/ML & Platform Operations)
Builds
Workday's digital experience, AI/ML platform, and Agent Factory initiative
Domain
Enterprise SaaS, AI/ML, Cloud Infrastructure
Required skills
LLM pipeline troubleshooting, root-cause analysis, cloud diagnostics (AWS/GCP/Azure), observability tooling (Grafana/Kibana), SQL, API debugging, Linux/Windows OS troubleshooting, incident management
Preferred skills
Explainable AI (XAI) concepts, severe incident de-escalation, cross-functional communication
Technologies
Kibana, Grafana, Datadog, Prometheus, AWS, GCP, Azure, Python, Bash, SQL, JSON, XML, Jira, ServiceNow, Salesforce
Responsibilities
Analyze system metrics and debug cloud-hosted ML service pipelines; Review LLM outputs and conversation logs to identify failure modes; Perform root-cause analysis on software defects and LLM agent execution failures; Triage high-severity support queues and manage critical customer escalations; Write complex SQL queries and scripts to validate backend data integrity; Partner with engineering teams to deliver permanent fixes
Seniority
Senior, hands-on IC
