CareerPlanSign in

Engineering Manager, Site Reliability Engineering

Foster City, CA💼 Full-time🗓 2026-09-25 → 2026-09-28

Core

Lead SRE across observability, incident management, load testing, performance engineering, cloud cost/capacity, and rollout infrastructure to support safe production changes and predictable performance for Replit's agentic software creation platform.

Role type

Engineering Manager, Site Reliability Engineering

Builds

Production platforms for observability, incident response, load testing, and performance engineering that enable safe software delivery and scaling.

Domain

Cloud infrastructure, distributed systems, SRE, AI-powered development platforms

Required skills

Engineering management, distributed systems architecture, Kubernetes, telemetry and observability, incident management, load testing, performance engineering, cloud cost optimization, team building and hiring, technical leadership

Preferred skills

GitOps, progressive-delivery platforms (Harness, ArgoCD, Kargo), OpenTelemetry, GCP cloud cost attribution, capacity planning, AI coding tools

Technologies

Kubernetes, OpenTelemetry, GCP, Harness, ArgoCD, Kargo

Responsibilities

Build and operate metrics, logs, traces, and alerting capabilities; own incident tooling and coordinate cross-team response; build and maintain load/failure testing capabilities; lead deep engagements on SLOs and end-to-end performance; review designs and debug difficult failure modes; coach engineers and develop technical leaders

Seniority

Manager, hands-on leadership

Sourced via ashby · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.