Senior Infrastructure Reliability Engineer
Core
Own the full lifecycle of critical self-hosted developer tools and on-prem compute platforms, ensuring high availability and reliability for the engineering organization.
Role type
Senior Infrastructure Reliability Engineer (SRE)
Builds
Core developer services (source control, CI/CD, artifact management) and on-prem GPU/simulation infrastructure
Domain
Defense technology, on-prem infrastructure, developer tooling
Deliverable
production ML models | infrastructure
Required skills
Linux administration, Kubernetes (bare-metal), Docker, Infrastructure-as-Code (Terraform), Configuration Management (Ansible/Puppet/Chef), Cloud platforms (AWS/GCP/Azure), Scripting (Python/Go/Bash), Incident response, SLO definition
Preferred skills
RKE2/k3s/kubeadm, Cilium, GitOps (ArgoCD/FluxCD), GPU/HPC workload schedulers (RunAI/Slurm/Kubeflow), JFrog Artifactory/Xray, CircleCI, GitHub Enterprise Server
Technologies
Terraform, Docker, Kubernetes, AWS, GCP, Azure, Ansible, Puppet, Chef, RHEL, Ubuntu, RunAI, JFrog Artifactory, CircleCI, GitHub Enterprise Server, Datadog, Prometheus, Grafana
Responsibilities
Operate and patch on-prem Kubernetes clusters and virtualization environments; Design and implement automated backup and upgrade systems; Scale infrastructure to support growing engineering workloads; Lead incident response and root cause analysis for critical services; Define and maintain SLOs for service availability; Manage CI/CD pipelines and developer tooling
Seniority
Senior, hands-on IC