Site Reliability Engineer
Core
Ensure the reliability, availability, and performance of a global infrastructure spanning 50+ data centers by improving observability, monitoring, alerting, and incident response.
Role type
Senior Site Reliability Engineer
Builds
Global infrastructure reliability and operational efficiency across compute, networking, storage, virtualization, and cloud environments
Domain
Cyber protection / Infrastructure / Cloud & On-prem
Deliverable
infrastructure
Required skills
Linux systems administration, Configuration management and automation (Ansible, Puppet, Terraform), Containerization and orchestration (Docker, Kubernetes), Monitoring and observability platforms (Prometheus, Grafana, ELK), Scripting/programming (Python, Go, Bash), Networking fundamentals (TCP/IP, DNS, load balancing)
Preferred skills
Bare-metal infrastructure management at scale, CI/CD systems (Jenkins, GitLab CI), Database administration (PostgreSQL, MySQL, Redis), Security best practices for production environments
Technologies
Ansible, Puppet, Terraform, Docker, Kubernetes, Prometheus, Grafana, ELK, Python, Go, Bash, Jenkins, GitLab CI, PostgreSQL, MySQL, Redis
Responsibilities
Design, implement, and maintain monitoring, alerting, and observability solutions; Respond to and troubleshoot production incidents to minimize downtime; Drive automation initiatives to reduce manual operational work; Optimize infrastructure across compute, networking, storage, virtualization, and cloud platforms; Collaborate with infrastructure and engineering teams to improve system resilience; Analyze metrics, logs, and system performance to proactively identify and resolve issues
Seniority
Senior, hands-on IC