Site Reliability Engineer — HPC & Automation (Silicon Engineering)
Core
Design, operate, scale, and automate high-performance computing infrastructure to accelerate chip design iterations, simulations, and regression for Starlink silicon development.
Role type
Site Reliability Engineer (HPC & Automation)
Builds
High-performance computing clusters and automated silicon simulation workflows for chip design teams.
Domain
Aerospace / Semiconductor / High-Performance Computing
Deliverable
infrastructure
Required skills
Linux system administration, Bash scripting, Python programming, HPC cluster management, CI/CD pipeline operations, performance bottleneck analysis
Preferred skills
Containerization (Docker, Kubernetes), Infrastructure as Code (Terraform, Ansible), Monitoring as Code (Grafana, Prometheus), ASIC design flow tools, Enterprise storage automation, REST API development
Technologies
Linux, Bash, Python, Docker, Kubernetes, Terraform, Ansible, Prometheus, Grafana, Jenkins, Slurm, LSF, MySQL, PostgreSQL, NetApp ONTAP, Cadence, Synopsys
Responsibilities
Deploy, upgrade, operate, maintain, and scale clusters and services; Develop automated turnkey solutions for silicon simulation workflows; Manage infrastructure as code and observability tools; Operate continuous integration pipelines; Identify and eliminate performance bottlenecks
Seniority
Individual Contributor