Tech Lead, Infrastructure Engineering (Cloud SRE L2)
Core
Design and operate reliable multi-cloud platforms and services, ensuring high availability, scalability, and performance through automation and incident response.
Role type
Senior IC Cloud Site Reliability Engineer (SRE)
Builds
Fault-tolerant, highly available cloud architectures and automated operational tooling
Domain
Fintech / Payments / Multi-cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Multi-cloud expertise (AWS/Azure/GCP), Infrastructure as Code (Terraform/Ansible), Container orchestration (Kubernetes/Docker), Observability (Splunk/Azure Monitor/Dynatrace), Python/PowerShell/Bash scripting, Incident response and root cause analysis, Capacity planning and scaling, Cloud data migration, Linux/Unix system administration
Preferred skills
Cloud certifications (AWS/Azure/GCP), Chaos engineering, Hybrid cloud deployments, SLO/SLI/Error budget management
Technologies
AWS, Azure, GCP, Terraform, Ansible, Splunk, Azure Monitor, Dynatrace, AWS CloudWatch, Kubernetes, Docker, Python, PowerShell, Bash
Responsibilities
Design and maintain fault-tolerant architectures across AWS, Azure, and GCP; Deploy and optimize cloud resources using IaC; Implement monitoring, alerting, and logging; Lead incident response and post-incident reviews; Drive capacity planning and scaling improvements; Build automation to reduce manual toil; Collaborate with Security on infrastructure hardening; Support cloud data migrations
Seniority
Senior, hands-on IC