Site Reliability Engineer
Core
Design, build, and maintain resilient cloud infrastructure solutions to support scalable and reliable applications, applying software engineering principles to operational challenges.
Role type
Senior Site Reliability Engineer (Cloud Infrastructure)
Builds
Cloud infrastructure, automation workflows, monitoring systems, and disaster recovery plans
Domain
Cloud Computing / Site Reliability Engineering
Deliverable
production ML models | infrastructure
Required skills
Cloud platform expertise (AWS/Azure/GCP), Containerization (Docker), Orchestration (Kubernetes), Scripting (Shell/Python), IaC (Terraform/CloudFormation), Configuration Management (Ansible/Puppet), Monitoring (Prometheus/Grafana/ELK), Observability (OpenTelemetry), Service Reliability Concepts (SLI/SLO/SLA)
Preferred skills
DevOps toolchain (Git/Jenkins/ArgoCD), Database management (MySQL/Hadoop), Cloud cost optimization, Gen AI exposure, Cloud security best practices, Disaster recovery planning, ITIL processes
Technologies
AWS, Azure, GCP, Docker, Kubernetes, Shell, Python, Ansible, Puppet, Terraform, CloudFormation, Prometheus, Grafana, ELK, OpenTelemetry, Rundeck, Jenkins, Git, ArgoCD, Crossplane, MySQL, Hadoop
Responsibilities
Design and manage cloud infrastructure for high availability and performance; Establish monitoring and alerting systems using SLI/SLO/SLA concepts; Automate infrastructure provisioning and repetitive tasks; Participate in incident response and root cause analysis; Collaborate with cross-functional teams on cloud solutions and POCs
Seniority
Senior, hands-on IC