Site Reliability Engineer (High Performance Computing)
Core
Operate and automate a shared High Performance Computing (HPC) platform supporting vehicle simulation, machine learning, and AI inference for rocket and satellite design.
Role type
Senior Site Reliability Engineer (HPC Infrastructure)
Builds
A shared compute platform for simulation, ML training, and AI inference workloads
Domain
Aerospace / High Performance Computing
Deliverable
production ML models | infrastructure
Required skills
Linux production operations, Infrastructure as Code, observability, automation scripting, capacity planning, Kubernetes administration, distributed storage management
Preferred skills
Prometheus/Grafana, Ansible/Terraform, Python, Docker/Podman, Slurm/PBS schedulers, CFD/FEA/ML workloads
Technologies
Linux, Kubernetes, Terraform, Ansible, Prometheus, Grafana, Python, Docker, Podman, Slurm, PBS, VAST
Responsibilities
Manage node lifecycle with infrastructure as code; build observability for cluster and job-level health; reduce toil with automation; lead capacity planning for compute and storage; participate in on-call rotation for incident response
Seniority
Senior, hands-on IC