SRE / DevOps Engineer
Core
Operate and continuously improve highly available production services and cloud infrastructure, leading incident management and automating software delivery.
Role type
Senior Site Reliability Engineer (SRE) / DevOps Engineer
Builds
Cloud-native microservices, CI/CD pipelines, and automated infrastructure on public clouds.
Domain
Cloud Infrastructure / DevOps / Site Reliability Engineering
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Kubernetes, Docker, CI/CD pipeline design, Linux administration, Infrastructure as Code (Terraform), Cloud platforms (AWS/Azure/GCP), Observability (Prometheus/Grafana/ELK), Scripting (Python/Bash/PowerShell), Incident management, Root cause analysis
Preferred skills
Service mesh (Istio), GitOps (ArgoCD), Message platforms (Kafka/RabbitMQ), Database administration, Cloud networking, Capacity planning, Java/Spring Boot, AIOps
Technologies
Kubernetes, Docker, Terraform, Ansible, Jenkins, GitLab CI, GitHub Actions, Azure DevOps, ArgoCD, Prometheus, Grafana, ELK, OpenTelemetry, Istio, Jaeger, Kafka, RabbitMQ, PostgreSQL, MongoDB, Redis, Python, Bash, Java, Spring Boot
Responsibilities
Operate and improve highly available production services; Lead incident management and root cause analysis; Design and maintain CI/CD pipelines; Deploy and operate containerized applications; Develop Infrastructure as Code and environment automation; Build monitoring, logging, and observability solutions; Automate operational tasks using scripting languages.
Seniority
Senior, hands-on IC