Observability Engineer - Infra
Core
Design, deploy, and maintain observability platforms and tools to provide comprehensive insights into system behavior, performance, and reliability for distributed systems.
Role type
Senior IC Observability Engineer (Infra)
Builds
Production observability solutions (metrics, logs, traces) and automation tooling for global distributed systems
Domain
Cloud-native infrastructure, DevOps, Observability
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
AWS, Kubernetes, Helm, Ansible, Python/Bash/Go, Docker, Terraform, OpenTelemetry, Prometheus, Grafana LGTM+, APM tools, eBPF, distributed systems architecture, networking, security principles
Preferred skills
Anomaly detection algorithms, predictive analytics, custom instrumentation library development, root cause analysis, mentorship
Technologies
Prometheus, Grafana, LGTM+, OpenTelemetry (OTel), Terraform, Ansible, Python, Bash, Go, Docker, Kubernetes, Helm, eBPF
Responsibilities
Design and maintain observability platforms; integrate observability into CI/CD pipelines; develop custom monitoring solutions and instrumentation; configure and optimize telemetry collection and storage; implement anomaly detection and predictive analytics; conduct root cause analysis; provide mentorship to junior team members; drive automation initiatives; participate in 24/7 on-call roster
Seniority
Senior, hands-on IC