Member of Technical Staff | Observability & Reliability
Core
Build, operate, and improve observability and reliability platforms for distributed systems across cloud and customer-hosted environments.
Role type
Senior IC platform engineering engineer (observability & reliability)
Builds
Observability platforms, telemetry pipelines, alerting systems, and reliability tooling for distributed workloads
Domain
Cloud infrastructure, distributed systems, Kubernetes, observability
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
OpenTelemetry, SLOs and error budgets, Kubernetes, Terraform, Helm, incident response, distributed systems debugging, production code development
Preferred skills
GCP/GKE or AWS/EKS experience, multi-node ML workloads, customer-hosted environment deployments, financial services domain knowledge
Technologies
OpenTelemetry, Kubernetes, Terraform, Helm, GCP, AWS, Ray
Responsibilities
Evolve and maintain the observability platform covering logs, metrics, traces, and alerting; Ensure environments report critical operational information to the control plane; Implement telemetry collection in customer Kubernetes environments; Detect and investigate infrastructure state discrepancies; Monitor health of deployment agents and ephemeral workloads; Define and maintain SLOs and actionable alerting; Lead incident response and postmortems; Optimize telemetry pipelines to reduce costs and improve signal quality; Write production-quality code and review technical changes