Senior Cloud, DevOps, Site Reliability Engineer (For Pooling)
Core
Designing, building, and supporting secure, scalable, and highly available cloud infrastructure and platform services in production environments.
Role type
Senior Cloud, DevOps, and Site Reliability Engineer (SRE)
Builds
Secure, scalable, and highly available cloud infrastructure and platform services
Domain
Cloud Engineering, DevOps, and Site Reliability Engineering
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
AWS production-grade cloud infrastructure, Terraform, Docker, Kubernetes, EKS, CI/CD tools (Jenkins, GitHub Actions, GitLab CI), monitoring and observability tools (Datadog, ELK, Splunk), scripting/automation (Python, Bash, Shell, PowerShell, Java, Groovy), Linux/Unix systems administration, networking, distributed systems, security fundamentals, incident response, troubleshooting, on-call responsibilities, reliability engineering concepts (monitoring, alerting, service health, recovery)
Preferred skills
high-availability, 24/7, transaction-heavy, customer-facing, or business-critical environments, fintech, regulated industries, or compliance-driven systems, Kafka (AWS MSK or Confluent), SQL and database performance analysis (Aurora PostgreSQL), hybrid or multi-cloud environments (AWS and Azure), AWS services (EC2, EKS, RDS/Aurora, Lambda, API Gateway, CloudFront, WAF, IAM), automation and configuration tools (Ansible, Chef), mentoring engineers
Technologies
AWS, Terraform, Docker, Kubernetes, EKS, Jenkins, GitHub Actions, GitLab CI, Datadog, ELK, Splunk, Python, Bash, Shell, PowerShell, Java, Groovy, Kafka, AWS MSK, Confluent, Aurora PostgreSQL, Ansible, Chef, EC2, Lambda, API Gateway, CloudFront, WAF, IAM
Responsibilities
Design, build, and support secure, scalable, and highly available cloud infrastructure and platform services; Develop and maintain Infrastructure as Code, automation scripts, and deployment pipelines; Build, manage, and optimize CI/CD pipelines; Support production environments through monitoring, alerting, observability, troubleshooting, and performance tuning; Participate in incident response, root cause analysis, post-incident reviews, and disaster recovery planning; Define, track, and improve SLIs, SLOs, SLAs, and other service reliability metrics; Create and maintain operational documentation such as runbooks, dashboards, and recovery procedures; Partner with software engineering, security, infrastructure, and product teams to improve platform reliability, operability, and delivery speed; Apply best practices in access control, security, compliance, availability, scalability, and cost optimization; Contribute to continuous improvement across tooling, standards, architecture, and operational processes; Provide technical leadership, mentorship, and guidance on cloud, DevOps, and SRE practices for senior-level opportunities
Seniority
Senior, hands-on IC with technical leadership and mentorship