Site Reliability Engineer
Core
Ensuring the availability, performance, security, and scalability of infrastructure and applications through automation and DevOps practices.
Role type
Senior Site Reliability Engineer (IC)
Builds
Scalable cloud infrastructure, CI/CD pipelines, and automated operational tooling
Domain
Cloud Infrastructure & DevOps
Deliverable
infrastructure
Required skills
Linux system administration, Infrastructure as Code (Terraform, Ansible, Chef), Container orchestration (Kubernetes), CI/CD pipeline management, Monitoring and logging (Datadog, Prometheus, Grafana), Scripting (Bash, Python, Go), Cloud platform management (AWS, Azure, GCP)
Preferred skills
Cloud certifications (AWS DevOps), Microservices architecture, Networking concepts (VPC, VPN, load balancing)
Technologies
Datadog, Prometheus, Grafana, ELK Stack, Chef, Terraform, Ansible, CloudFormation, Jenkins, GitLab CI, CircleCI, Docker, Kubernetes, AWS, Azure, GCP, Bash, Python, Go, Git
Responsibilities
Develop and enhance monitoring systems and lead incident response for production outages; Design and maintain scalable infrastructure using IaC tools; Ensure stability and performance of Linux-based infrastructure; Build and manage CI/CD pipelines for automated deployments; Develop scripts and tooling to automate repetitive operational tasks; Collaborate with security teams to ensure compliance with security standards; Identify and resolve performance bottlenecks in systems and applications
Seniority
Senior, hands-on IC