Site Reliability Engineer - Product Support
Core
Building observability, monitoring, and operational foundations for PTaaS and vulnerability management products.
Role type
Junior-to-Mid SRE (Product Support)
Builds
Observability stacks, incident response playbooks, and operational automation for security products.
Domain
Cybersecurity / SaaS Platform Engineering
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Python scripting, Linux command line, cloud compute fundamentals, incident response, runbook authoring, alert triage, CI/CD pipeline support, Docker, REST APIs, Git
Preferred skills
Kubernetes (GKE), OpenTelemetry, vector databases, LLM integration, Terraform, Checkly, Prometheus, Grafana Cloud
Technologies
Python, Bash, GCP, AWS, Azure, Docker, Kubernetes, GKE, Grafana, Datadog, CloudWatch, Loki, ELK, GitHub Actions, Cloud Build, OpenTelemetry, Django, Flask, Node.js, Checkly, Prometheus, RAG, ChromaDB, Pinecone, OpenAI, Anthropic
Responsibilities
Monitor and triage alerts from Grafana Cloud, Checkly, and load balancer logs; Participate in on-call rotation as a supporting responder; Assist in maintaining availability and performance dashboards; Monitor scan job success rates and PTaaS report delivery timelines; Automate recurring operational tasks using Python scripts and Bash; Assist with OpenTelemetry instrumentation tasks on Python and Node.js services; Support deployment validation and rollback procedures in CI/CD pipelines; Contribute to writing and maintaining runbooks and operational playbooks.
Seniority
Junior/Mid, hands-on IC with mentorship