Site Reliability / Operations Engineer (TS/SCI)
Core
Develop, maintain, and troubleshoot automated monitoring, metrics data collection, and real-time dashboards within a mission-focused environment.
Role type
Junior to mid-level Site Reliability / Operations Engineer
Builds
Automated monitoring solutions, metrics data pipelines, and operational dashboards
Domain
National security / Systems engineering / Observability
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Linux system administration, Bash scripting, ELK Stack (Elasticsearch, Logstash, Kibana), system monitoring, log aggregation, operational analytics, networking fundamentals, troubleshooting methodologies
Preferred skills
Python scripting, cloud or hybrid environment support, CI/CD concepts, Agile development
Technologies
Linux (RHEL, CentOS, Ubuntu), Bash, Elasticsearch, Logstash, Kibana
Responsibilities
Develop and maintain automation solutions in Linux environments; Create and enhance Bash scripts for system administration and log gathering; Design and maintain dashboards using Kibana; Support Elasticsearch data ingestion and search optimization; Monitor system performance and identify automation opportunities; Collaborate with stakeholders to implement solutions; Assist with log analysis and operational reporting; Document automation processes and configurations; Participate in system testing and deployment activities.
Seniority
Junior to mid-level, hands-on IC
