Staff Site Reliability Engineer
Core
Own internal systems infrastructure for an AI robotics company developing autonomous humanoid robots.
Role type
Staff Site Reliability Engineer
Builds
Cloud and on-prem infrastructure enabling critical operations like CI/CD, software distribution, and manufacturing systems.
Domain
Robotics / Cloud Infrastructure
Deliverable
infrastructure
Required skills
Linux/Unix systems administration, programming/scripting, cloud platforms (Azure, AWS, GCP), high-availability distributed systems design, infrastructure as code (Terraform, CloudFormation, Ansible), monitoring and logging tools (Prometheus, Grafana, Datadog), networking fundamentals, SLO definition and incident response management.
Preferred skills
Migrate SaaS to self-hosted solutions, reduce human workload through automation, data-driven optimization.
Technologies
Azure, AWS, GCP, Terraform, CloudFormation, Ansible, Prometheus, Grafana, Datadog
Responsibilities
Manage mission critical infrastructure including Source Configuration Management, CI/CD, and software distribution; Migrate SaaS to self-hosted solutions; Implement monitoring, alerting, and incident response plans; Automate deployment and scaling; Define SLOs and manage systems assets; Partner with security team on remediations.
Seniority
Staff, hands-on IC with strategic scope