Site Reliability Engineer
Core
Designing, deploying, and managing enterprise cloud solutions with a focus on high availability, fault tolerance, and cost optimization.
Role type
Site Reliability Engineer (Cloud Infrastructure)
Builds
Scalable cloud infrastructure, CI/CD pipelines, and deployment automation.
Domain
Financial services / Cloud Infrastructure
Deliverable
infrastructure
Required skills
Cloud platform expertise (AWS, GCP, Azure), Infrastructure-as-Code (Terraform, CloudFormation), Scripting/Programming (Python, Go, Bash), Networking fundamentals, Containerization and orchestration (Docker, Kubernetes), Observability tools (Datadog, Prometheus, Grafana), Linux/Unix administration
Preferred skills
Financial services industry experience, Compliance frameworks (SOC 2, ISO 27001), Cloud certifications
Technologies
AWS, GCP, Azure, Terraform, CloudFormation, Bicep, Docker, Kubernetes, Datadog, Prometheus, Grafana
Responsibilities
Design and maintain scalable cloud infrastructure; Collaborate with software engineers to embed reliability best practices; Lead incident response and conduct post-mortems; Develop and maintain CI/CD pipelines; Optimize cloud resource utilization and cost management; Ensure adherence to financial industry security and compliance standards; Document systems and processes
Seniority
Mid-level (3–5 years experience)