Site Reliability Engineer, Kubernetes Platform (Starshield)
Core
Design, operate, and scale on-premise Kubernetes infrastructure and automation for the world's largest government satellite constellation.
Role type
Site Reliability Engineer, Kubernetes Platform
Builds
On-premise compute resources, databases, monitoring, and distributed storage for Starshield
Domain
National security satellite communications / Cloud Infrastructure
Deliverable
infrastructure
Required skills
Kubernetes, Terraform, Ansible, Linux, Bash, Python, C++, Go, TCP/IP
Preferred skills
Kubernetes cluster management, Linux boot process, CI/CD, Bazel, distributed databases, performance optimization
Technologies
Kubernetes, Terraform, Ansible, Bazel, Makefiles
Responsibilities
Develop automation to deploy and manage on-premise Kubernetes clusters; Deploy and manage core infrastructure such as databases, monitoring and distributed storage; Collaborate with software engineers to create scalable products; Engage in the whole lifecycle of services from inception to refinement; Monitor and alert supporting systems for high availability; Perform hands-on integration and troubleshooting across the Starshield stack; Identify areas for improvement to enable high system availability
Seniority
Mid-level, hands-on IC