CareerPlanSign in

DevOps / Site Reliability Engineer

Singapore, Singapore💼 Full-time🗓 2026-06-11 → 2026-09-26

Core

Build and maintain the core infrastructure of the AIOps platform, including unified monitoring, alerting, and FinOps cost observability systems for a leading crypto exchange.

Role type

DevOps / Site Reliability Engineer (AIOps & Observability)

Builds

AIOps platform, monitoring & alerting systems, FinOps cost observability platform, internal R&D infrastructure

Domain

Cryptocurrency / Blockchain / Cloud Infrastructure

Deliverable

production ML models | infrastructure

Required skills

Python, Go or Java, major cloud platforms (Alibaba Cloud or AWS), monitoring stacks (Prometheus, Grafana, ELK), CI/CD toolchains (GitLab CI, Nexus), container registries, cloud security operations

Preferred skills

AIOps or observability platform development, full-stack capability (React/Vue + backend), AI/LLM application development (LLM API, RAG, Agent frameworks)

Technologies

Python, Go, Java, React, Vue, Alibaba Cloud, AWS, CloudWatch, CloudMonitor, Prometheus, Grafana, ELK, GitLab CI, Nexus

Responsibilities

Build and maintain the core infrastructure of the AIOps platform; Maintain and optimize internal R&D infrastructure; Manage monitoring data collection, alert governance, and cost data visualization; Support cloud security operations and compliance auditing

Seniority

Mid-level (3+ years experience)

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.