Site Reliability Engineer - Saint Louis
Core
Operate and maintain developer tooling infrastructure (GitLab, Jira, Confluence, PostgreSQL) and ensure platform reliability, availability, and security for mission-critical engineering environments.
Role type
Site Reliability Engineer (Infrastructure & Developer Tooling)
Builds
Developer platforms, CI/CD pipelines, and secure engineering environments
Domain
Defense/Aerospace, Cloud Infrastructure, DevOps
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Linux system administration, Git-based source control workflows, cloud computing (AWS/Azure/GCP), Infrastructure as Code (Ansible), incident response, root cause analysis, backup and restore procedures, scripting (Bash/Python/PowerShell), SQL troubleshooting, change management
Preferred skills
Active security clearance, Agile software development, Jira/Confluence administration, PostgreSQL administration, monitoring and alerting tools, disaster recovery planning, technical investigation leadership
Technologies
AWS, Ansible, Atlassian, Azure, Bash, C#, Confluence, DevSecOps, DevOps, Git, GitLab, JIRA, Java, Jenkins, Linux, PostgreSQL, PowerShell, Python, SQL, SAP
Responsibilities
Administer, maintain, and upgrade developer tools (GitLab, Jira, Confluence, SonarQube); develop and maintain Infrastructure as Code and automation scripts; plan and execute approved changes including patches and upgrades; support incident response and post-incident reviews; monitor system health and operational metrics; triage and resolve service requests and pipeline issues; maintain runbooks and disaster recovery procedures; partner with developers and cybersecurity personnel to improve platform reliability.
Seniority
Associate through Senior (Level 3/4 preferred 5+ or 9+ years experience)