Technical Lead, Platform Engineering (Observability)
Core
Build and operate an end-to-end observability platform (metrics, logs, traces, alerting) for 500+ microservices to ensure system reliability and reduce incident detection/mitigation times.
Role type
Senior Platform Engineer (Observability)
Builds
AI-powered incident detection/response tools, self-service observability tooling, and standardized monitoring/alerting infrastructure.
Domain
Cloud Infrastructure / Observability / SRE
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Go or Python, Terraform, GCP/AWS, distributed tracing, alerting system design, SLO/SLI frameworks, data pipeline design, internal tool development
Preferred skills
AI for observability (anomaly detection, alert correlation), MTTD/MTTM improvement in microservices
Technologies
Datadog, Prometheus, Grafana, Kubernetes, Terraform, GCP, AWS
Responsibilities
Design and operate the observability stack; drive improvements in MTTD and MTTM; build AI-powered anomaly detection and alert correlation; develop self-service instrumentation and monitoring tooling; define observability standards and SLO frameworks; automate operational workflows; mentor team members and lead technical decisions.
Seniority
Senior, hands-on IC with mentorship responsibilities