CareerPlanSign in

Site Reliability Engineer (High Performance Computing)

Hawthorne, CA💼 Full-time💰 $1–$1🗓 2026-09-23 → 2026-09-26

Core

Operate and automate a shared High Performance Computing (HPC) platform supporting vehicle simulation, machine learning, and AI inference for rocket and satellite design.

Role type

Senior Site Reliability Engineer (HPC Infrastructure)

Builds

A shared compute platform for simulation, ML training, and AI inference workloads

Domain

Aerospace / High Performance Computing

Deliverable

production ML models | infrastructure

Required skills

Linux production operations, Infrastructure as Code, observability, automation scripting, capacity planning, Kubernetes administration, distributed storage management

Preferred skills

Prometheus/Grafana, Ansible/Terraform, Python, Docker/Podman, Slurm/PBS schedulers, CFD/FEA/ML workloads

Technologies

Linux, Kubernetes, Terraform, Ansible, Prometheus, Grafana, Python, Docker, Podman, Slurm, PBS, VAST

Responsibilities

Manage node lifecycle with infrastructure as code; build observability for cluster and job-level health; reduce toil with automation; lead capacity planning for compute and storage; participate in on-call rotation for incident response

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.