CareerPlanSign in

Senior Infrastructure Reliability Engineer

Costa Mesa, CA (OC-00)💼 Full-time💰 $166,000–$166,000🗓 2026-09-04 → 2026-09-27

Core

Own the full lifecycle of critical self-hosted developer tools and on-prem compute platforms, ensuring high availability and reliability for the engineering organization.

Role type

Senior Infrastructure Reliability Engineer (SRE)

Builds

Core developer services (source control, CI/CD, artifact management) and on-prem GPU/simulation infrastructure

Domain

Defense technology, on-prem infrastructure, developer tooling

Deliverable

production ML models | infrastructure

Required skills

Linux administration, Kubernetes (bare-metal), Docker, Infrastructure-as-Code (Terraform), Configuration Management (Ansible/Puppet/Chef), Cloud platforms (AWS/GCP/Azure), Scripting (Python/Go/Bash), Incident response, SLO definition

Preferred skills

RKE2/k3s/kubeadm, Cilium, GitOps (ArgoCD/FluxCD), GPU/HPC workload schedulers (RunAI/Slurm/Kubeflow), JFrog Artifactory/Xray, CircleCI, GitHub Enterprise Server

Technologies

Terraform, Docker, Kubernetes, AWS, GCP, Azure, Ansible, Puppet, Chef, RHEL, Ubuntu, RunAI, JFrog Artifactory, CircleCI, GitHub Enterprise Server, Datadog, Prometheus, Grafana

Responsibilities

Operate and patch on-prem Kubernetes clusters and virtualization environments; Design and implement automated backup and upgrade systems; Scale infrastructure to support growing engineering workloads; Lead incident response and root cause analysis for critical services; Define and maintain SLOs for service availability; Manage CI/CD pipelines and developer tooling

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 814,000+ jobs from 20+ sources.