Site Reliability Engineer - I
Core
Ensures production services remain available, scalable, and efficient on Google Cloud Platform by bridging development and operations for containerized infrastructure.
Role type
Entry-level Site Reliability Engineer (SRE)
Builds
Containerized applications and infrastructure on Google Cloud Platform
Domain
Cloud infrastructure, networking, and observability
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
GitOps (ArgoCD), Linux internals, TCP/IP networking, Kubernetes (GKE), Python/Bash/Go scripting, Grafana ecosystem, alert triage, incident response
Preferred skills
AIOps concepts, anomaly detection, log-based ML models
Technologies
ArgoCD, GKE, Grafana, Mimir, Prometheus, Loki, Tempo, tcpdump, curl, dig, traceroute
Responsibilities
Deploy and manage containerized application lifecycles via GitOps pipelines; investigate and resolve infrastructure, OS, and network alerts; diagnose connectivity and latency issues across cloud VPCs and Kubernetes overlays; troubleshoot OS bottlenecks (CPU, memory, storage); utilize AIOps tools for event correlation and root-cause analysis; participate in on-call rotations to mitigate production issues
Seniority
Junior, hands-on IC