[MLA] Senior Site Reliability Engineer (SRE) – Kubernetes
Core
Own production reliability for ServiceNow's AI Experience Framework stack, including an SSR runtime (karuna) built on Lit and server-rendered web components, paired with a ServiceNow Glide/Java platform layer.
Role type
Senior Site Reliability Engineer (Kubernetes)
Builds
AI-first user interfaces and SSR runtime platform
Domain
Enterprise SaaS / AI Platform / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes production operations, incident response and root cause analysis, Linux and networking fundamentals, Node.js and JVM/Java troubleshooting, CI/CD and GitOps, observability (Splunk, Prometheus, Grafana), service-to-service authentication (mTLS, JWT)
Preferred skills
Web Components/Lit debugging, canary rollout operations, distributed tracing, event-driven autoscaling (KEDA), enterprise platform integration
Technologies
Kubernetes, Lit, Node.js, Java/JVM, Splunk, Prometheus, Grafana, Helm, ArgoCD, Flux, HTTP/2
Responsibilities
Support deployment, operation, and reliability of production services on Kubernetes; monitor service health and investigate incidents; participate in on-call support, incident response, and postmortems; troubleshoot application runtime and networking issues; support CI/CD and observability pipelines
Seniority
Senior, hands-on IC
