Site Reliability Engineer
Core
Guardian of production uptime and system stability for cloud services, bridging development and operations through observability and incident response.
Role type
Site Reliability Engineer (SRE)
Builds
Production cloud services and containerized environments
Domain
Cloud Infrastructure / DevOps
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Azure cloud management, Kubernetes/Docker orchestration, Infrastructure as Code (Terraform/Ansible), scripting (Bash/PowerShell), metrics and logging systems, Root Cause Analysis, CI/CD pipeline support
Preferred skills
.NET/ASP.NET environment knowledge, KQL, relational databases (MSSQL/PostgreSQL)
Technologies
Azure, AWS, GCP, Prometheus, Grafana, Datadog, Azure App Insights, Terraform, Ansible, Chef, Docker, Kubernetes, Logstash, Dynatrace, KQL, MSSQL, PostgreSQL
Responsibilities
Design and maintain monitoring dashboards and alerting systems; lead troubleshooting and perform Root Cause Analysis for production incidents; manage technical escalations and communicate with clients; support deployment via CI/CD pipelines; automate operational tasks using scripting and IaC; monitor KPIs and optimize containerized environments
Seniority
Mid-level, hands-on IC