CareerPlanSign in

Sr. Staff Lead Site Reliability Engineer (R5803)

San Mateo, California💼 Full-time🗓 2026-09-03 → 2026-09-26

Core

Establish and mature SRE practices for cloud infrastructure and platform services, ensuring predictable operation and recovery.

Role type

Sr. Staff Lead Site Reliability Engineer (IC + Team Lead)

Builds

Cloud infrastructure and platform services for defense-tech autonomy systems

Domain

Defense technology, cloud infrastructure, distributed systems

Deliverable

production ML models | infrastructure

Required skills

SRE practices, SLIs/SLOs definition, observability (monitoring/alerting/logging/tracing), incident response & root-cause analysis, AWS cloud operations, infrastructure-as-code, containerized applications, Python/Go automation, distributed systems debugging, technical roadmap planning

Preferred skills

SRE function establishment, Kubernetes, regulated environment operations, capacity planning, cloud cost management, cross-organization infrastructure support

Technologies

AWS, Python, Go, Kubernetes

Responsibilities

Define and implement SLIs/SLOs; Build monitoring/alerting/logging/tracing; Lead complex incident response and RCA; Identify and eliminate recurring failure modes; Improve resilience via automation/testing/capacity planning; Develop operational tooling; Partner on reliability requirements in system design; Mentor engineers on SRE practices; Define and manage SRE roadmap

Seniority

Sr. Staff, hands-on IC with team leadership

Sourced via lever · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.