CareerPlanSign in

Member of Technical Staff | Observability & Reliability

US🌐 Remote💼 Full-time🗓 2026-09-23 → 2026-09-26

Core

Build, operate, and improve observability and reliability platforms for distributed systems across cloud and customer-hosted environments.

Role type

Senior IC platform engineering engineer (observability & reliability)

Builds

Observability platforms, telemetry pipelines, alerting systems, and reliability tooling for distributed workloads

Domain

Cloud infrastructure, distributed systems, Kubernetes, observability

Deliverable

production ML models | product features | dashboards & analysis | infrastructure

Required skills

OpenTelemetry, SLOs and error budgets, Kubernetes, Terraform, Helm, incident response, distributed systems debugging, production code development

Preferred skills

GCP/GKE or AWS/EKS experience, multi-node ML workloads, customer-hosted environment deployments, financial services domain knowledge

Technologies

OpenTelemetry, Kubernetes, Terraform, Helm, GCP, AWS, Ray

Responsibilities

Evolve and maintain the observability platform covering logs, metrics, traces, and alerting; Ensure environments report critical operational information to the control plane; Implement telemetry collection in customer Kubernetes environments; Detect and investigate infrastructure state discrepancies; Monitor health of deployment agents and ephemeral workloads; Define and maintain SLOs and actionable alerting; Lead incident response and postmortems; Optimize telemetry pipelines to reduce costs and improve signal quality; Write production-quality code and review technical changes

Sourced via lever · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.