云监控/APM研发工程师-火山引擎
Core
Design and develop a large-scale monitoring platform and APM system for internal business and To B customers, focusing on high-throughput, low-latency time-series data infrastructure and intelligent monitoring models.
Role type
Senior IC backend engineer (monitoring & APM)
Builds
High-throughput, low-latency monitoring data infrastructure and intelligent monitoring models (AIOps)
Domain
Cloud infrastructure, observability, distributed systems
Deliverable
production ML models
Required skills
Go/Java/Python, Linux, TCP/IP, HTTP, gRPC/Thrift/bRPC, Redis, MySQL, Kafka, RocketMQ, high-concurrency distributed system design, SLO definition, alerting optimization
Preferred skills
Prometheus/OpenFalcon/Zabbix source code, VictoriaMetrics/InfluxDB/OpenTSDB, APM/distributed tracing (OpenTelemetry/SkyWalking/Jaeger), high-availability stability (disaster recovery, rate limiting, self-healing), AIOps