Lead SRE- Observability
Core
Design and operate scalable observability and telemetry platforms to enable engineering teams to monitor, troubleshoot, and run distributed services reliably at scale.
Role type
Lead Site Reliability Engineer (Observability)
Builds
Scalable observability and telemetry platforms processing logs, metrics, traces, and events for cloud infrastructure and engineering organizations.
Domain
Healthcare technology, Cloud Infrastructure, Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Linux systems engineering, Cloud infrastructure, SRE practices, Observability platform design, Infrastructure as Code, Automation engineering, Python, Golang, Bash, Distributed systems troubleshooting, Cloud-native environments (AWS), Containerized platforms, Monitoring strategy, Telemetry pipelines, Incident response, Root cause analysis, Operational excellence
Preferred skills
High-scale telemetry/analytics platforms, Kubernetes, Docker, CI/CD systems, Networking troubleshooting (tcpdump, Wireshark), Cross-functional leadership, Healthcare technology experience
Technologies
OpenSearch/Elasticsearch, Kafka, Prometheus, Grafana, Vector, Fluentd, OpenTelemetry, ClickHouse, Terraform, CloudFormation, AWS, Linux
Responsibilities
Build and operate scalable observability and telemetry platforms; Support monitoring, alerting, and instrumentation strategies; Design resilient, automated infrastructure and platform services; Troubleshoot complex production issues involving distributed systems, Linux infrastructure, networking, cloud services, and telemetry pipelines; Participate in incident response and on-call processes; Mentor engineers on SRE best practices, observability strategy, and scalable systems design
Seniority
Senior, hands-on IC with technical leadership