Senior Site Reliability Engineer
Core
Design, build, and operate shared platform foundations (GCP, Kubernetes, networking, CI/CD, observability) for a high-scale AI-powered content operating system serving global customers.
Role type
Senior Site Reliability Engineer (IC)
Builds
Scalable, high-availability cloud infrastructure and content distribution systems for enterprise clients.
Domain
Cloud Infrastructure / Content Operations / AI
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, GCP, CI/CD pipeline design, observability stack (Prometheus), distributed systems debugging, CDN/edge/caching architecture, incident response, on-call management, automation scripting, code review, technical mentorship
Preferred skills
Experience with high request volume systems, modernizing edge layers, building golden paths for deployments
Technologies
Kubernetes, Prometheus, ElasticSearch, PostgreSQL, NATS, Kong, Fastly, Google Cloud Platform
Responsibilities
Design and operate GCP infrastructure, Kubernetes, networking, routing, CI/CD, and observability; Diagnose and troubleshoot complex distributed systems; Ensure observability and analyze stack behavior; Contribute to modernizing edge, caching, and gateway layers; Raise reliability bar through dashboards, alert standards, and incident response; Build automation for safe rollouts and production readiness; Mentor engineers through code and design reviews; Participate in on-call rotation.
Seniority
Senior, hands-on IC
