Sr. Manager, Site Reliability Engineer
Core
Lead a global Site Reliability Engineering (SRE) organization to advance a modern operating model combining centralized reliability capabilities with product-aligned support.
Role type
Senior Manager, Site Reliability Engineering
Builds
Global SRE organization, hybrid reliability operating model, enterprise observability standards, cloud financial governance
Domain
Enterprise SaaS (Hiring Platform), Cloud Infrastructure, Site Reliability Engineering
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Site Reliability Engineering, AWS cloud architecture, observability, cloud engineering, automation, incident management, problem management, FinOps, technical leadership, distributed team management
Preferred skills
Infrastructure as Code (Terraform), cloud cost optimization, service blueprinting, architectural evaluation
Technologies
AWS (ECS, ECR, EC2, RDS, Aurora, S3, DynamoDB, OpenSearch, SQS, SNS, Kinesis, IAM, Organizations), Grafana, OpenTelemetry, Sumo Logic, New Relic, CloudWatch, Terraform
Responsibilities
Define and execute global SRE strategy and operating model; establish hybrid SRE model with centralized and product-aligned support; develop engineers and technical leaders through coaching and mentorship; partner with Product Engineering to establish SLIs, SLOs, and error budgets; lead incident response and root-cause analysis; drive enterprise observability strategy and governance; incorporate cloud financial awareness into architecture and engineering decisions; review and guide complex AWS architectures for resiliency and cost efficiency
Seniority
Senior, hands-on IC with leadership responsibilities