Principal Site Reliability Engineer ( U.S Citizenship required )
Core
Architect and operate high-scale, self-healing cloud infrastructure for end-to-end digital experience monitoring and security services.
Role type
Principal Site Reliability Engineer (Cloud Infrastructure)
Builds
Production-grade synthetic and Real User Monitoring (RUM) platforms on GCP/AWS
Domain
Cybersecurity / Cloud Infrastructure / Observability
Deliverable
production ML models | infrastructure
Required skills
Kubernetes (GKE/EKS), Infrastructure as Code (Terraform, Helm, Ansible), Cloud-native architecture (GCP/AWS), Python, Go, Linux internals, GitOps (GitLab CI, ArgoCD), Root Cause Analysis (RCA), Policy-as-code automation
Preferred skills
Data streaming frameworks (Kafka, Apache Pulsar), AI productivity tools (Claude, Cursor, Windsurf, GitHub Copilot), FedRAMP/SOC2 compliance frameworks
Technologies
Terraform, Kubernetes, GitLab CI, ArgoCD, Prometheus, Grafana, Loki, Docker, GCP, AWS, Vault, Kafka, MySQL, Python, Bash, Go
Responsibilities
Architect "Golden Paths" for service delivery integrating SLOs and automated canary analysis; Design and operate reliable, secure Cloud infrastructure for high-scale monitoring; Develop automation frameworks for Infrastructure as Code and Monitoring as Code; Lead root cause analysis of critical production issues; Drive CI/CD and AIOps initiatives for self-healing infrastructure
Seniority
Principal, hands-on IC with strategic architecture