Site Reliability Engineer, GNC
Core
Operate and scale custom-built, mission-critical products and infrastructure for the Guidance, Navigation, and Control (GNC) teams to support vehicle design, trajectory optimization, and high-fidelity simulations for reusable launch systems.
Role type
Site Reliability Engineer (Infrastructure & HPC)
Builds
On-prem services, large-scale Monte Carlo simulations on HPC clusters, automated data analysis pipelines, and continuous integration systems for rocket and simulation software.
Domain
Aerospace / High-Performance Computing / Guidance, Navigation, and Control
Deliverable
production ML models | infrastructure
Required skills
Linux administration, Python development, HPC cluster management, incident response, performance optimization, configuration management, virtualization, networking (TCP/IP), database understanding.
Preferred skills
Docker, Kubernetes, Ansible, Terraform, build systems (Bazel, Make), GPU fleet management, large-scale data analysis systems.
Responsibilities
Deploy, upgrade, operate, and scale mission-critical GNC products; provision and maintain virtual and physical servers; monitor and maintain HPC clusters; collaborate with software engineers to create maintainable products; manage computational infrastructure; optimize application performance; provide end-user support for analysis applications.
Seniority
Mid-level IC (2+ years experience)