Site Reliability Engineer (SRE)
Core
Design, implement, and maintain scalable, reliable infrastructure; automate deployment, scaling, and management of applications and services.
Role type
Site Reliability Engineer (SRE)
Builds
Cloud infrastructure, automated deployment pipelines, and monitoring systems
Domain
Cloud infrastructure and DevOps
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Linux/Unix system administration, Python, Bash, Go, Terraform, CloudFormation, Kubernetes, Docker, Prometheus, Grafana, Ansible, Jenkins, GitLab CI, capacity planning, incident management
Preferred skills
AWS Certified Solutions Architect, GCP Professional Cloud Architect
Technologies
AWS, GCP, Azure, Kubernetes, Docker, Terraform, CloudFormation, Prometheus, Grafana, Nagios, Datadog, Ansible, Chef, Puppet, Jenkins, GitLab CI, CircleCI, ELK Stack
Responsibilities
Design and maintain scalable infrastructure, automate deployment and scaling, monitor system health and troubleshoot issues, participate in on-call rotations, develop runbooks and automation scripts, collaborate with development teams on architecture, conduct performance tuning and capacity planning, improve observability, document operational procedures and incident post-mortems
Seniority
Mid-Senior, hands-on IC