Staff Site Reliability Engineer (m/f/d)
Core
Transform central Cloud & IT services into highly reliable, observable, and automated products for engineering teams.
Role type
Staff Site Reliability Engineer
Builds
Shared platform services including Vault/PKI, CI/CD systems, monitoring platforms, and self-hosted tools.
Domain
Defense technology / Cloud Infrastructure
Deliverable
infrastructure
Required skills
Site Reliability Engineering, DevOps, Platform Engineering, production system operations, secrets management, observability principles, automation scripting (Python, Go, shell), incident response, runbook creation.
Preferred skills
personal project automation, homelab experience, proactive problem-solving.
Technologies
Vault, PKI, CI/CD, Python, Go, shell
Responsibilities
Define SLOs and incident response workflows, build observability practices with metrics and dashboards, create resilient deployment and recovery patterns, automate recurring operational tasks, partner with engineering teams on service ownership, collaborate on backend application operability, participate in incident response and post-mortems.
Seniority
Staff, hands-on IC with strategic impact