CareerPlanSign in

后端开发工程师(监控平台) - DCS

北京💼 Full-time🗓 2026-09-28

Core

Design and develop a unified monitoring and observability platform for physical machines, cloud hosts, GPUs, and IaaS resources to support high availability and stability for ByteDance's global data center operations.

Role type

Senior backend engineer (monitoring platform)

Builds

Unified monitoring architecture, data pipelines, alerting engines, and disaster recovery systems for data center infrastructure.

Domain

Cloud infrastructure / Data center operations

Deliverable

production ML models | product features

Required skills

Golang, C/C++, system architecture design, time-series storage, alerting engine development, high-throughput low-latency optimization, LLM integration for fault prediction

Preferred skills

Deep understanding of open-source monitoring systems (Prometheus, VictoriaMetrics), source code modification, automated root cause analysis

Technologies

Golang, C, C++, Prometheus, VictoriaMetrics, LLM

Responsibilities

Design system monitoring platform architecture; Build data collection, transmission, storage, and query pipelines; Implement disaster recovery and high availability mechanisms; Explore intelligent monitoring using LLMs for fault prediction and diagnosis.

Sourced via bytedance · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.