Senior Site Reliability Engineer
Core
Building, scaling, and maintaining mission-critical infrastructure for spacecraft and company-wide systems, including cloud services and embedded software.
Role type
Senior Site Reliability Engineer (IC)
Builds
Cloud-based services, embedded software for spacecraft, and company-wide operational infrastructure.
Domain
Aerospace / Space Infrastructure / Cloud Systems
Deliverable
production ML models | product features | infrastructure
Required skills
Kubernetes, Terraform, Prometheus, Grafana, InfluxDB, Python, Bash, PowerShell, Linux, software-defined networking
Preferred skills
Azure cloud infrastructure, GitOps (ArgoCD), containerd, Docker, HPC environments (Slurm), hybrid cloud/on-prem, distributed systems debugging, database modeling
Technologies
Kubernetes, Terraform, Prometheus, Grafana, InfluxDB, Azure, Ansible, Salt, ArgoCD, containerd, Docker, Slurm
Responsibilities
Deploy, maintain, and operate mission-critical applications and infrastructure; Build and evolve Infrastructure as Code (IaC) frameworks; Implement and operate observability systems; Build and maintain CI/CD pipelines; Partner with software and hardware engineers to deliver reliable systems; Identify, analyze, and resolve system bottlenecks and reliability risks; Respond to and resolve production incidents; Rotate through the team's on-call schedule.
Seniority
Senior, hands-on IC