Staff Site Reliability Engineer
Core
Own and evolve the engineering-wide SRE strategy, operating model, and reliability standards for a cybersecurity platform (NodeZero) serving IT Ops/SecOps teams, consultants, and MSSPs.
Role type
Staff Site Reliability Engineer (Strategy & Standards)
Builds
Production-safe autonomous pentests and assessment operations on AWS and Kubernetes
Domain
Cybersecurity / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Distributed systems design, Reliability engineering, Observability, Incident management, Backend automation, Infrastructure as Code, CI/CD pipelines, AWS, Kubernetes
Preferred skills
Cross-functional leadership, Technical direction setting, On-call model design
Technologies
Python, Terraform, Datadog, New Relic, Grafana, Gitlab CI, ArgoCD
Responsibilities
Define organization-wide SLIs, SLOs, and error budgets; Lead cross-functional alignment on reliability and incident response; Establish observability standards and actionable alerting; Drive complex cross-functional reliability initiatives; Shape the technical direction and growth path of the SRE function; Participate in 24/7 on-call rotation.
Seniority
Staff, hands-on IC with strategic scope