Senior Site Reliability Engineer
Core
Operate, improve, and sustain mission-critical developer platforms and infrastructure for Air Dominance engineering teams.
Role type
Senior Site Reliability Engineer (IC)
Builds
Developer tooling infrastructure, CI/CD pipelines, cloud/on-prem infrastructure, and automated operational workflows.
Domain
Aerospace / DevOps & Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
GitLab, CI/CD, AWS, Azure, Linux, PostgreSQL, Infrastructure as Code, Ansible, Kubernetes, Docker, technical leadership, incident response, automation, monitoring, SLOs/SLIs
Preferred skills
GitLab runners administration, Jira/Confluence administration, Artifactory, SonarQube, disaster recovery planning, classified/air-gapped environment experience, Security+ certification
Technologies
GitLab, Jira, Confluence, PostgreSQL, AWS, Azure, Ansible, Docker, Kubernetes, Artifactory, SonarQube
Responsibilities
Operate and maintain developer tooling infrastructure; lead deployment and configuration of build servers and CI/CD pipelines; administer cloud and on-premises infrastructure; serve as technical owner for platform reliability and performance; troubleshoot complex infrastructure and application issues; develop and maintain Infrastructure as Code and automation; plan and execute approved changes and patches; establish and improve monitoring, alerting, and operational metrics; lead software development tool administration and integration; support incident response and root cause analysis; mentor junior engineers; improve runbooks and disaster recovery procedures; evaluate platform risks and recommend improvements; present technical status to management.
Seniority
Senior, hands-on IC with technical leadership