Engineering Manager, Site Reliability Engineering
Core
Lead SRE across observability, incident management, load testing, performance engineering, cloud cost/capacity, and rollout infrastructure to support safe production changes and predictable performance for Replit's agentic software creation platform.
Role type
Engineering Manager, Site Reliability Engineering
Builds
Production platforms for observability, incident response, load testing, and performance engineering that enable safe software delivery and scaling.
Domain
Cloud infrastructure, distributed systems, SRE, AI-powered development platforms
Required skills
Engineering management, distributed systems architecture, Kubernetes, telemetry and observability, incident management, load testing, performance engineering, cloud cost optimization, team building and hiring, technical leadership
Preferred skills
GitOps, progressive-delivery platforms (Harness, ArgoCD, Kargo), OpenTelemetry, GCP cloud cost attribution, capacity planning, AI coding tools
Technologies
Kubernetes, OpenTelemetry, GCP, Harness, ArgoCD, Kargo
Responsibilities
Build and operate metrics, logs, traces, and alerting capabilities; own incident tooling and coordinate cross-team response; build and maintain load/failure testing capabilities; lead deep engagements on SLOs and end-to-end performance; review designs and debug difficult failure modes; coach engineers and develop technical leaders
Seniority
Manager, hands-on leadership