Senior Site Reliability Engineer (SRE) – CloudVision as a Service (CVaaS)
Core
Build and operate the global CloudVision-as-a-Service (CVaaS) fleet, an enterprise network management and streaming telemetry SaaS offering, ensuring scalability, reliability, and stability.
Role type
Senior Site Reliability Engineer (SRE)
Builds
CloudVision SaaS platform (Kubernetes-native) and supporting data platforms (NetDL)
Domain
Cloud networking, SaaS, streaming telemetry, Kubernetes
Deliverable
production ML models | product features | infrastructure
Required skills
Distributed systems architecture, Kubernetes operations, Python, Golang, Bash scripting, CI/CD pipeline management, disaster recovery planning, observability stack management, capacity planning, database management (distributed/managed), cloud-native security
Preferred skills
GCP/GKE experience, Ansible/Pulumi
Technologies
Kubernetes, GCP, GKE, Golang, Python, Ansible, Pulumi, Bash
Responsibilities
Drive architecture and performance projects for the Data Platform (NetDL), lead capacity planning and autoscaling initiatives, manage disaster recovery and observability strategies, oversee change management and CI/CD processes, optimize service network architecture and cloud costs, ensure cloud-first application security, lead sustainable incident response and blameless postmortems
Seniority
Senior, hands-on IC