Sr. Site Reliability Engineer - Top Secret Clearance (Starlink)
Core
Design, build, test, and operate distributed systems and infrastructure for Starlink, the world's largest satellite constellation, to provide global broadband internet.
Role type
Senior Site Reliability Engineer (Infrastructure & Distributed Systems)
Builds
Sharded, geo-redundant distributed systems; multi-region deployment and monitoring infrastructure; petabyte-scale bare metal compute clusters.
Domain
Aerospace / Satellite Communications / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Linux, distributed systems, sharding, geo-redundancy, multi-region architecture, petabyte-scale compute management, software development lifecycle, performance optimization, version control, continuous integration, testing, hardware troubleshooting, network-layer troubleshooting, Python, C#, Java, Scala, Go
Preferred skills
Kubernetes, Istio, in-stream data processing, Apache Kafka, Spark, HBase, HDFS, Flink
Responsibilities
Upgrade existing distributed systems to become sharded and geo-redundant in multiple data centers; Advance existing deployment, monitoring, and alerting infrastructure to support a multi-region environment; Manage petabyte scale bare metal compute clusters; Closely collaborate with engineers across all programs to create highly operable, scalable, and maintainable products; Engage throughout the whole software development lifecycle of services -- from inception to design, deployment, operation, and iterative refinement; Focus on performance bottlenecks and performance improvement techniques.
Seniority
Senior, hands-on IC