Site Reliability Engineer (Manufacturing Infrastructure)
Core
Ensure reliability, stability, and scalability of compute, storage, and networking infrastructure supporting SpaceX manufacturing systems for Starship, Starlink, Starshield, and Terafab.
Role type
Site Reliability Engineer (Manufacturing Infrastructure)
Builds
Factory uptime, throughput, and production scale for aerospace hardware programs
Domain
Aerospace manufacturing infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Linux system administration, software development fundamentals, infrastructure as code, containerization and virtualization, database management, capacity planning, incident response, observability
Preferred skills
Production compute/storage/networking experience, Terraform/Ansible/Puppet, Docker/Kubernetes/vSphere/QEMU/KVM, Postgres/Clickhouse, first-principles problem solving, multi-site operations
Technologies
Linux, Terraform, Ansible, Puppet, Docker, Kubernetes, vSphere, QEMU, KVM, Postgres, Clickhouse
Responsibilities
Deploy, upgrade, operate, maintain, and scale compute, storage, and networking for manufacturing systems; Manage infrastructure as code and use observability to monitor platform health; Design for reliability, stability, and scale; Practice proactive maintenance including capacity planning and lifecycle management; Partner with software engineers and manufacturing stakeholders to build operable systems; Conduct sustainable incident response and blameless postmortems; Provide high-quality support to manufacturing and engineering users; Participate in on-call rotations and travel to sites for deployments and incidents
Seniority
Mid-level, hands-on IC
