Site Reliability Engineer (SRE) Manager
Core
Lead a Network Operations Center (NOC) L2 team responsible for triaging, investigating, and resolving escalated network and platform incidents for a global smart home service platform.
Role type
Senior Engineering Manager (Site Reliability / Network Operations)
Builds
Incident response processes, runbooks, and operational systems for a SaaS platform serving ISPs and smart home devices.
Domain
Smart home technology, ISP services, network operations, cloud infrastructure.
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
People leadership, incident management, root-cause analysis, networking fundamentals, observability tooling, process improvement, stakeholder communication, team scheduling, hiring and coaching.
Preferred skills
Telecom/ISP industry experience, cloud infrastructure (AWS/GCP/Azure), container orchestration (Kubernetes/Docker), automation scripting (Python/Bash), scaling NOC teams, ITIL certification.
Technologies
Grafana, Prometheus, Datadog, PagerDuty, Splunk, AWS, GCP, Azure, Kubernetes, Docker, Python, Bash.
Responsibilities
Lead and mentor NOC L2 engineers; own on-call rotation and escalation policies; drive root-cause analysis and post-incident reviews; partner with Engineering/Infrastructure/Product to reduce incidents; establish and refine operational runbooks; monitor operational metrics (MTTD, MTTR); manage team schedules and coverage; hire and develop engineers; act as escalation point for critical incidents; collaborate with L1 leadership; drive blameless postmortem culture.
Seniority
Senior, hands-on IC with people management