Site Reliability Engineer
Required skills
Linux/Unix systems administration, Source control tools (Git preferred), Configuration Management and Infrastructure as Code tools (Ansible, Puppet, Terraform preferred), Container technology (Docker, Kubernetes preferred), Monitoring tools (Prometheus, Grafana, Nagios, or similar.), Bash scripting, Python scripting, Incident management, Troubleshooting, Root cause analysis, Postmortems, Incident response plans, Real-world build systems (Jenkins, DroneCI, or similar), Systems Architecture, Systems Design, Implementation, Maintenance, Operation, HTTP Service APIs, Virtualization (VMWare, Proxmox, Oracle Linux Virtualization Manager), Network administration, Security and Testing frameworks, Compliant regulated industries (Finance, Healthcare, Government), Distributed data processing, Databases, Large-scale file systems.
Preferred skills
Bachelor's degree in information systems, computer science, technology, or a related field, 2+ years of relevant and/or equivalent experience (in lieu of degree), Experience with non-cloud infrastructure, Experience running a large-scale 24/7 production environment, Experience with distributed data processing, databases, and large-scale file systems.
Responsibilities
Design and implement solutions to reduce toil, Ensure reliability of critical products and services, Instantiate and maintain production infrastructure using Infrastructure as Code and Configuration Management tools, Build and maintain proper monitoring of services using centralized logging and time series databases, Automate deployments, administration, and monitoring of services following CI/CD practices, Work with engineering and information security teams to enhance, document, establish processes and improve operability and security, Participate in team on-call rotation.
Seniority
Mid-level (3+ years of software and/or operational experience required).
Domain
Site Reliability Engineering, Certificate Lifecycle Management (CLM), Digital Trust, Cybersecurity.