Kubernetes Site Reliability Engineer
Core
Managing and sustaining on-premises and cloud-based Kubernetes clusters to support space launch telemetry analysis, modeling, simulation, and AI/LLM training/inference services for national space assets.
Role type
Senior Site Reliability Engineer (Kubernetes Platform)
Builds
Production Kubernetes clusters and PaaS services for space engineering teams
Domain
Aerospace / Space Systems / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes cluster management, Linux systems administration, cloud infrastructure (AWS/Azure), security patching, capacity planning, performance tuning, automation scripting, networking fundamentals, storage fundamentals
Preferred skills
Rancher, persistent container storage (Portworx/Rook/OpenEBS/Longhorn), Linux performance tuning, security hardening, Infrastructure-as-Code (Terraform/Ansible), GitOps (ArgoCD), Golang/Python
Responsibilities
Develop and sustain advanced services for Kubernetes-based PaaS, manage production Kubernetes environments on-prem and cloud, perform full security patching and backups, resolve engineering problems independently, provide after-hours support during launch events, support scientists running applications, evaluate and test new technologies
Seniority
Senior, hands-on IC with architecture responsibilities