DevOps Engineer (Observability)
Core
Re-architecting Twilio's observability platform to unify telemetry flows (logs, metrics, traces) into a scalable, OpenTelemetry-first stack for faster incident response and cost-effective operations.
Role type
Senior IC Platform Engineering Engineer (Observability)
Builds
Unified observability platform components, developer tooling, and APIs for incident response and performance analysis.
Domain
Cloud Infrastructure / Observability / Distributed Systems
Deliverable
production ML models | product features | infrastructure
Required skills
Building and scaling observability systems, OpenTelemetry, distributed systems architecture, high-scale telemetry design, AWS, Kubernetes, infrastructure-as-code, Go/Python/Java
Preferred skills
ClickHouse, Grafana Mimir, Athena, FinOps tooling, open-source observability contributions
Technologies
OpenTelemetry, S3, ClickHouse, Kafka, Prometheus, AWS, Kubernetes, Go, Python, Java
Responsibilities
Lead end-to-end architecture and delivery of observability platform components; Drive consistency across logs, metrics, traces, and profiling; Serve as technical advisor and mentor across the platform org; Go deep in specific problem areas like high-cardinality telemetry; Collaborate with product teams and SREs to integrate observability into workflows; Design developer-friendly tooling and APIs.
Seniority
Senior, hands-on IC with architectural leadership