Sr. Staff Site Reliability Engineer (Linux/Network troubleshooting/Scripting)
Core
Architecting, scaling, and maintaining next-generation cloud infrastructure for a Zero Trust Exchange platform, bridging development and operations to ensure high availability and security.
Role type
Sr. Staff Site Reliability Engineer (Cloud Infrastructure)
Builds
Cloud-native Zero Trust Exchange platform, containerized architectures, observability systems
Domain
Cybersecurity / Cloud Infrastructure
Deliverable
production ML models | product features | infrastructure
Required skills
Linux/BSD system administration, Kubernetes (EKS/GKE), Terraform, Ansible, Python, Go, C, Java, CI/CD pipeline design, web protocols (HTTP, SSL/TLS, DNS), observability architecture (Grafana, SLIs/SLOs)
Preferred skills
OS/software packaging and distribution, incident prevention process design, mentoring
Technologies
EKS, GKE, Grafana, Terraform, Ansible, AWS, Kubernetes
Responsibilities
Design and implement cloud management automations to eliminate toil; oversee and optimize containerized architectures; lead creation and optimization of scalable monitoring and alerting systems; own cloud operations, deployments, on-call support, and incident management; serve as a core member of cross-functional project teams consulting on concept feasibility
Seniority
Sr. Staff, hands-on IC with strategic impact