Systems Engineering, Metrics and Alerting
Core
Design, deliver, and operate a scalable observability platform (metrics, alerting, error tracking, logging, tracing) for internal engineering teams.
Role type
Senior Systems Engineer (Observability Platform)
Builds
Internal observability stack and metrics/alerting pipelines
Domain
Internet infrastructure, distributed systems, observability
Deliverable
production ML models | product features | infrastructure
Required skills
Go, distributed Linux environments, high-scale distributed system design, TSDBs/Columnar stores, Prometheus, Alertmanager, Thanos, networking protocols (OSI L2-7, BGP)
Preferred skills
high-bandwidth transit internetworking, code simplicity and performance optimization
Responsibilities
Solve scaling bottlenecks in the Metrics & Alerting pipeline, operate highly distributed and scalable systems, participate in global on-call rotation, research and introduce cutting-edge technologies, contribute to open-source
Seniority
Senior, hands-on IC