Site Reliability Engineer
Core
Building the reliability engineering function from the ground up for a complex, distributed SaaS platform that manages and secures global networks.
Role type
Foundational Site Reliability Engineer (SRE)
Builds
A network digital twin and SaaS platform for enterprise customers
Domain
Network security, cloud infrastructure, and SaaS reliability
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
SRE practices definition (SLOs, SLIs, error budgets), observability infrastructure design, incident response leadership, SDLC reliability integration, Kubernetes orchestration, network fundamentals (TCP/IP, DNS, routing), cloud platform expertise (AWS/GCP/Azure), infrastructure as code, Python/Bash scripting
Preferred skills
Experience with network management or observability platforms, supporting enterprise/federal government customers, early-stage SRE hire experience
Technologies
Kubernetes, Prometheus, Grafana, Datadog, Splunk, Terraform, Ansible, AWS, GCP, Azure
Responsibilities
Define and drive SRE practices from the ground up, drive reliability and operational excellence of the SaaS platform, build and maintain observability infrastructure, lead incident response and post-mortems, partner with engineering to embed reliability thinking into the SDLC, help define and build the SRE team as the company scales
Seniority
Senior, foundational hire with path to leadership