Sr. Staff Lead Site Reliability Engineer (R5803)
Core
Establish and mature SRE practices for cloud infrastructure and platform services, ensuring predictable operation and recovery.
Role type
Sr. Staff Lead Site Reliability Engineer (IC + Team Lead)
Builds
Cloud infrastructure and platform services for defense-tech autonomy systems
Domain
Defense technology, cloud infrastructure, distributed systems
Deliverable
production ML models | infrastructure
Required skills
SRE practices, SLIs/SLOs definition, observability (monitoring/alerting/logging/tracing), incident response & root-cause analysis, AWS cloud operations, infrastructure-as-code, containerized applications, Python/Go automation, distributed systems debugging, technical roadmap planning
Preferred skills
SRE function establishment, Kubernetes, regulated environment operations, capacity planning, cloud cost management, cross-organization infrastructure support
Technologies
AWS, Python, Go, Kubernetes
Responsibilities
Define and implement SLIs/SLOs; Build monitoring/alerting/logging/tracing; Lead complex incident response and RCA; Identify and eliminate recurring failure modes; Improve resilience via automation/testing/capacity planning; Develop operational tooling; Partner on reliability requirements in system design; Mentor engineers on SRE practices; Define and manage SRE roadmap
Seniority
Sr. Staff, hands-on IC with team leadership