Site Reliability Engineer
Core
Operate and improve reliability of large-scale hybrid infrastructure (on-premises colocation and multi-cloud Azure/AWS/GCP) supporting enterprise middleware, databases, and AI workloads.
Role type
Junior/Entry-level Site Reliability Engineer
Builds
Reliable, secure, and performant platform infrastructure for an investment bank
Domain
Financial Services / Platform Engineering / Hybrid Cloud Infrastructure
Deliverable
infrastructure
Required skills
Linux/Windows server administration, networking fundamentals (TCP/IP, DNS, DHCP, VLANs, firewalls), scripting (Python, Bash, or PowerShell), cloud computing concepts (Azure, AWS, or GCP), incident response, observability
Preferred skills
Infrastructure-as-Code (Terraform, Ansible), GitOps, containerization (Docker, Kubernetes), observability tools (Datadog, Prometheus, Grafana), cloud security posture management
Technologies
Azure, AWS, GCP, Terraform, Ansible, Docker, Kubernetes, Datadog, Prometheus, Grafana, Python, Bash, PowerShell
Responsibilities
Monitor production systems and respond to incidents, maintain on-premises and cloud infrastructure, support multi-cloud operations, contribute to automation and IaC, participate in DevSecOps/GitOps workflows, support cloud security remediation, participate in on-call rotation
Seniority
Entry-level, hands-on IC