Site Reliability Engineer
Core
Own the reliability, health, and deployment automation of the defense AI platform, monitoring dashboards and managing incident response.
Role type
Senior Site Reliability Engineer
Builds
Secure, robust, high-availability delivery pipeline and observability tooling
Domain
Defense technology / AI systems / Cloud infrastructure
Deliverable
infrastructure
Required skills
incident response, debugging, observability tooling, capacity planning, deployment automation, workflow optimization, self-service tool development, secure pipeline implementation
Preferred skills
SLO/SLI definition, CI/CD process design, containerization, web services architecture, relational database design, Elasticsearch/OpenSearch
Technologies
Datadog, AWS, Terraform, Pulumi, Python, Bash, Docker
Responsibilities
Monitor system telemetry to detect health issues and performance degradation; own end-to-end debugging and incident response; build logging and observability tooling; develop self-service automation tools; improve engineering workflows and deployment pipelines; communicate system status and post-incident learnings
Seniority
Senior, hands-on IC