SRE / Platform engineer
Core
Own and evolve the systems that keep infrastructure fast, reliable, and developer-friendly, leading observability, CI/CD, and GitOps initiatives.
Role type
Senior Site Reliability Engineer / Platform Engineer
Builds
Production-grade observability stacks, automated CI/CD pipelines, and scalable Kubernetes environments.
Domain
Cloud Infrastructure, DevOps, Data Engineering
Deliverable
production ML models | infrastructure
Required skills
Observability stack management (Prometheus, Thanos, VictoriaMetrics, Grafana, PromQL), Kubernetes administration (EKS), GitOps implementation (ArgoCD, Flux), CI/CD pipeline design, Infrastructure as Code (Terraform, Pulumi), Scripting (Python, Bash, Go), Stateful service management (MySQL, ClickHouse, Kafka), Security best practices in CI/CD.
Preferred skills
Experience with data lake architecture, Real-time data systems, Cost optimization strategies.
Technologies
EKS, FastAPI, MySQL, ClickHouse, Kafka, Prometheus, Thanos, VictoriaMetrics, Grafana, ArgoCD, Flux, Terraform, Pulumi, Python, Bash, Go.
Responsibilities
Own and evolve the observability stack from data collection to dashboarding; Implement GitOps workflows to automate application lifecycles; Maintain and scale EKS clusters for high availability and cost efficiency; Collaborate on data lake architecture and tooling; Define and enforce security best practices within CI/CD and AWS environments; Continuously improve monitoring, alerting, and incident response playbooks.
Seniority
Senior, hands-on IC