Site Reliability Engineer
Core
Design, deploy, and operate enterprise-grade cloud applications ensuring reliability, availability, and performance through automation and incident management.
Role type
Senior Site Reliability Engineer (SRE) / DevOps
Builds
Production cloud applications on Google Cloud Platform (GCP) using Kubernetes
Domain
Financial Services / Cloud Infrastructure
Deliverable
production ML models | product features | infrastructure
Required skills
Kubernetes administration, Terraform, CI/CD pipeline design, Helm charts, service mesh configuration, scripting (Java/Python/Go/Bash), cloud platform expertise
Preferred skills
Anthos Service Mesh, Openshift Cloud, GCP Secret Manager, Prometheus/Grafana
Technologies
Google Kubernetes Engine (GKE), Terraform, GitHub Actions, Helm, Istio, Anthos Service Mesh, Docker, Prometheus, Grafana, GCP
Responsibilities
Implement monitoring and alerting systems, manage Kubernetes infrastructure including node management and auto-scaling, develop automation tools for deployment and scaling, respond to system outages and conduct post-mortem analysis, optimize application and infrastructure performance, design and manage CI/CD pipelines, configure service networking components
Seniority
Senior, hands-on IC