Member of Technical Staff | Observability & Reliability
Core
Platform team expert responsible for evolving observability stacks, ensuring telemetry reporting across cloud and on-premise dataplanes, and managing incident response to maintain 99.9% serving availability.
Role type
Senior IC Platform Engineer (Observability & Reliability)
Builds
Observability stack (logs, metrics, traces, alerting) and telemetry agents for cloud and customer-hosted environments
Domain
Cloud infrastructure, distributed systems, financial services/regulated environments
Deliverable
production ML models | infrastructure
Required skills
OpenTelemetry, observability backends, SLOs, error budgets, incident management, Kubernetes, Infrastructure as Code (Terraform/Helm), production code quality
Preferred skills
GCP/GKE, AWS/EKS, ML multi-node/multi-cluster workloads, shipping to customer-hosted Kubernetes
Technologies
OpenTelemetry, Kubernetes, Terraform, Helm, Ray
Responsibilities
Evolve observability stack for logs, metrics, traces, and alerting; Ensure dataplanes report health, heartbeat, logs, metrics, and usage; Detect drift between desired and actual state; Monitor deployment and runtime agent health; Define SLOs and lead incident response/postmortems; Reduce telemetry costs
Seniority
Senior, hands-on IC