Softwareentwickler/in
Core
Design, implement, and maintain highly reliable and scalable infrastructure and services using cloud platforms, focusing on automation, monitoring, and incident management.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Cloud infrastructure and services on GCP
Domain
Cloud Infrastructure / DevOps
Deliverable
production ML models | infrastructure
Required skills
GCP, Terraform, Ansible, SaltStack, Docker, Kubernetes, Prometheus, Grafana, Java, Python, Go, Bash, Linux, Networking (DNS, TCP/IP, HTTP), Security (SSL/TLS, IAM)
Preferred skills
CI/CD pipelines, Incident Management, Capacity planning, Root cause analysis
Technologies
GCP, Terraform, Ansible, SaltStack, Gitlab, Prometheus, Grafana, ELK Stack, Fluentd, Splunk, Jira, ServiceNow, Git, SVN
Responsibilities
Automate repetitive tasks using infrastructure-as-code tools; Establish and monitor SLIs/SLOs; Perform capacity planning and optimization; Lead incident management and post-mortem processes; Collaborate with development and operations teams for smooth deployment.
Seniority
Senior, hands-on IC