Sr. Kubernetes Platform Site Reliability Engineer (Starlink)
Core
Design, operate, and scale on-premise Kubernetes clusters and core infrastructure (databases, monitoring, distributed storage) to support Starlink's global satellite constellation and network.
Role type
Senior Site Reliability Engineer (Kubernetes Platform)
Builds
On-premise compute resources, highly scalable software products, and the infrastructure running the world's largest satellite constellation.
Domain
Aerospace / Satellite Communications / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes cluster management, Linux system administration, Infrastructure as Code (Terraform, Ansible), scripting (Bash, Python), software development (Python, C++, Go), distributed storage, monitoring and alerting, TCP/IP networking.
Preferred skills
Python development frameworks, Linux boot process knowledge, CI/CD pipelines, build technologies (Bazel, Makefiles), performance optimization, distributed database modeling, large-scale server management.
Responsibilities
Develop automation to deploy and manage on-premise Kubernetes clusters; Deploy and manage core infrastructure such as databases, monitoring and distributed storage; Collaborate with software engineers to create scalable products; Engage in the full lifecycle of services from design to operation; Monitor and alert on systems to ensure high availability; Perform hands-on integration and troubleshooting across the Starlink stack; Identify areas for improvement and create innovative solutions for high system availability.
Seniority
Senior, hands-on IC