Senior Site Reliability Engineer (m/f/x)
Core
Design, implement, and maintain highly reliable and scalable infrastructure and services using cloud platforms, ensuring system reliability and performance.
Role type
Senior Site Reliability Engineer (IC)
Builds
Scalable and efficient cloud infrastructure and services
Domain
Cloud Infrastructure / Site Reliability Engineering
Deliverable
production ML models | infrastructure
Required skills
Cloud infrastructure (GCP), Containerization (Docker, Kubernetes), Scripting/Programming (Java, Python, Go, Bash), Monitoring/Observability (Prometheus, Grafana, ELK, Fluentd, Splunk), Networking (DNS, TCP/IP, HTTP), Security (SSL/TLS, firewalls, IAM), System Administration (Linux, Windows), CI/CD pipelines, Incident Management, Capacity Planning, Automation (Terraform, Ansible, SaltStack)
Preferred skills
Experience with Jira, ServiceNow, Git, SVN
Technologies
GCP, Terraform, Ansible, SaltStack, Gitlab, Prometheus, Grafana, Docker, Kubernetes, Java, Python, Go, Bash, ELK Stack, Fluentd, Splunk, Jira, ServiceNow, Git, SVN
Responsibilities
Design and maintain scalable cloud infrastructure; Automate repetitive tasks; Collaborate on deployment and operation; Establish and monitor SLIs/SLOs; Perform capacity planning; Lead incident management and post-mortems
Seniority
Senior, hands-on IC