Site Reliability Engineer (West Coast)
Core
Design and implement solutions to reduce toil and ensure reliability of critical services and infrastructure.
Role type
Site Reliability Engineer
Builds
Production infrastructure, monitoring systems, and automated deployment pipelines.
Domain
Cybersecurity / Certificate Lifecycle Management (CLM)
Deliverable
infrastructure
Required skills
Linux/Unix systems administration, Infrastructure as Code, Configuration Management, CI/CD practices, centralized logging, time series databases, container technology, incident management, root cause analysis, postmortems, incident response planning, scripting (Bash, Python), HTTP Service APIs, virtualization, network administration, distributed data processing, databases, large-scale file systems
Preferred skills
Source control tools (Git), monitoring tools (Prometheus, Grafana, Nagios), non-cloud infrastructure, real-world build systems (Jenkins, DroneCI), Security and Testing frameworks, compliant regulated industries experience (Finance, Healthcare, Government)
Technologies
Ansible, Puppet, Terraform, Docker, Kubernetes, Prometheus, Grafana, Nagios, Jenkins, DroneCI, VMWare, Proxmox, Oracle Linux Virtualization Manager
Responsibilities
Ensure reliability of critical products and services by meeting SRE objectives; Instantiate and maintain production infrastructure using IaC and Configuration Management; Build and maintain proper monitoring utilizing centralized logging and time series databases; Automate deployments, administration, and monitoring following CI/CD practices; Work with engineering and security teams to enhance processes and improve operability and security; Participate in team on-call rotation.
Seniority
Mid-level, hands-on IC