Sr. Site Reliability Engineer (Starlink)
Core
Upgrade distributed systems to be sharded and geo-redundant, manage petabyte-scale bare metal clusters, and advance deployment/monitoring infrastructure for Starlink's global satellite internet.
Role type
Senior Site Reliability Engineer (Infrastructure & Distributed Systems)
Builds
Starlink satellite internet system (hardware, receivers, and software)
Domain
Aerospace / Satellite Communications / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Linux, Kubernetes, Istio, Python, C#, Java, Scala, Go, version control, CI/CD, monitoring, performance optimization, data processing (Kafka, Spark, HBase, HDFS, Flink), hardware troubleshooting
Preferred skills
5+ years SRE/DevOps experience, on-premise Kubernetes/Istio deployment, in-stream data processing, network-layer troubleshooting
Technologies
Kubernetes, Istio, Apache Kafka, Spark, HBase, HDFS, Flink, Linux
Responsibilities
Upgrade distributed systems to be sharded and geo-redundant across multiple data centers; Manage petabyte-scale bare metal compute clusters; Focus on performance bottlenecks and improvement techniques; Collaborate with engineers to create scalable and maintainable products; Engage throughout the software development lifecycle from inception to iterative refinement
Seniority
Senior, hands-on IC