Kubernetes Site Reliability Engineer
Core
Develop and maintain on-premises and cloud-based Kubernetes clusters forming a PaaS for space launch telemetry analysis, modeling, simulation, and AI/LLM training/inference.
Role type
Senior Site Reliability Engineer (Kubernetes Platform)
Builds
Production Kubernetes clusters, Coder workspaces, Kueue, Knative, Crossplane control planes
Domain
Aerospace / Space Systems / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes cluster management, Linux systems administration, cloud infrastructure (AWS/Azure), security patching, capacity planning, performance tuning, automation scripting, networking fundamentals, storage fundamentals
Preferred skills
Rancher, persistent container storage (Portworx/Rook/Longhorn), Linux performance tuning, IaC/GitOps (ArgoCD/Terraform/Ansible), Golang/Python
Responsibilities
Manage Kubernetes clusters for production on-prem and cloud environments; perform full security patching and maintain high uptime; resolve full Kubernetes stack engineering problems independently; support real-time analysis of space launch telemetry data; provide after-hours support during launch events; evaluate and test new technologies; automate operations with code
Seniority
Senior, hands-on IC
