Site Reliability Engineer
Core
Design, build, and maintain resilient cloud infrastructure solutions to support scalable and reliable applications, ensuring operational resilience and continuous stability.
Role type
Senior Site Reliability Engineer (IC)
Builds
Cloud infrastructure, automation workflows, and monitoring systems for high availability and performance.
Domain
Cloud Infrastructure & Site Reliability Engineering
Deliverable
infrastructure
Required skills
Cloud platforms (AWS/Azure/GCP), Containerization (Docker), Orchestration (Kubernetes), Scripting (Shell/Python), Infrastructure as Code (Terraform/CloudFormation), Monitoring (Prometheus/Grafana/ELK), OpenTelemetry, Linux internals, Virtualization (VMware/Hyper-V/KVM), Networking protocols
Preferred skills
DevOps toolchain (Git/Jenkins/ArgoCD), Database management (MySQL/Hadoop), Cloud cost optimization, Gen AI, Cloud security best practices, Disaster recovery planning, ITIL processes
Technologies
AWS, Azure, GCP, Docker, Kubernetes, Python, Shell, Ansible, Puppet, Terraform, CloudFormation, Prometheus, Grafana, ELK, OpenTelemetry, Rundeck, Jenkins, VMware, Hyper-V, KVM, MySQL, Hadoop, Git, ArgoCD, Crossplane