[MLA] Senior Site Reliability Engineer - AI Experience Framework
Core
Own production reliability for ServiceNow's AI Experience Framework stack, including an SSR runtime (karuna) and a Glide/Java platform layer.
Role type
Senior Site Reliability Engineer (AI Experience Framework)
Builds
SSR runtime (karuna), Glide/Java platform layer, multi-tier proxy/HTTP2 routing chain, sharded V8 isolate pools
Domain
Enterprise SaaS, AI-first user interfaces, Server-Side Rendering
Deliverable
production ML models | infrastructure
Required skills
Kubernetes operations, Node.js troubleshooting (V8 isolates, event loops), JVM troubleshooting (GC, thread dumps), Prometheus alerting, Grafana dashboards, Linux/networking fundamentals, CI/CD (Helm, GitOps), Splunk, in-memory caching (Valkey/Redis), mTLS management, incident response
Preferred skills
Server-side rendering architectures, multi-version/canary rollout strategies, ServiceNow Glide platform, distributed tracing, event-driven autoscaling (KEDA)
Responsibilities
Own Kubernetes deployment and operational health, build and maintain production observability, diagnose and resolve Node.js and JVM production incidents, own incident response (runbooks, on-call, postmortems), drive CI/CD and infrastructure-as-code, partner on reliability gaps and capacity planning, represent reliability concerns in architecture reviews
Seniority
Senior, hands-on IC
