Sr. High Performance Computing (HPC) Systems Engineer
Core
Administer and manage HPC clusters, storage systems, and high-speed networks to support SpaceX's engineering disciplines and proprietary systems.
Role type
Senior IC HPC Systems Engineer
Builds
Linux-based compute clusters, storage systems, and high-speed networks for SpaceX engineering teams
Domain
Aerospace / High Performance Computing
Deliverable
production ML models | infrastructure
Required skills
Linux systems administration, Kubernetes, enterprise networking, virtualization, security technologies, cluster resource managers (Slurm, PBS, LSF), monitoring and alerting (Prometheus, Grafana, Nagios), configuration management (Puppet, Ansible), scripting (Bash, Python)
Preferred skills
HPC cluster deployment and troubleshooting, scientific and engineering computing (CFD, FEA), large scale AI training, GPU usage and CUDA, containerization (Docker, Podman, Singularity)
Responsibilities
Administer and manage HPC clusters, storage systems, and high-speed networks; Provide application support to SpaceX employees; Install and integrate Linux-based compute clusters; Write instructional documentation and convey highly technical ideas in non-technical terms
Seniority
Senior, hands-on IC