LI Site Reliability Engineer SB25
Core
Design, build, and maintain Linux-based observability platforms including logging, metrics, and tracing systems for large-scale deployments.
Role type
Site Reliability Engineer (Infrastructure)
Builds
Production observability platforms (logging, metrics, tracing) and supporting visualization tools
Domain
Infrastructure Engineering / Observability / Linux Systems
Deliverable
production ML models | product features | infrastructure
Required skills
Linux system administration, Bash scripting, Kubernetes, containerization, network packet capture, server hardware troubleshooting, TCP/IP and OSI model knowledge, automation tooling
Preferred skills
Ruby, Puppet, Ansible, Terraform, kernel optimization, CPU optimization, GitOps methodology
Technologies
Elasticsearch, Logstash, Kibana, OpenSearch, Graylog, MongoDB, Grafana, Loki, Prometheus, Mimir, Metrictank, Tempo, Jaeger, Kafka, RedPanda, Redis, Min.io, nProbe, Plixar Scrutinizer, ntopng, Arkime, Elastic Beats, Fluentbit, OpenTelemetry
Responsibilities
Design and maintain logging, metrics, and tracing platforms; support network capture visibility platforms and packet brokers; troubleshoot server hardware and Linux systems; research and document system usage; optimize kernel and CPU performance for network captures
Seniority
Mid-level to Senior, hands-on IC