CareerPlanSign in

Member of Technical Staff | Observability & Reliability

São Paulo💼 Full-time🗓 2026-09-23 → 2026-09-26

Core

Platform team expert responsible for evolving observability stacks, ensuring telemetry reporting across cloud and on-premise dataplanes, and managing incident response to maintain 99.9% serving availability.

Role type

Senior IC Platform Engineer (Observability & Reliability)

Builds

Observability stack (logs, metrics, traces, alerting) and telemetry agents for cloud and customer-hosted environments

Domain

Cloud infrastructure, distributed systems, financial services/regulated environments

Deliverable

production ML models | infrastructure

Required skills

OpenTelemetry, observability backends, SLOs, error budgets, incident management, Kubernetes, Infrastructure as Code (Terraform/Helm), production code quality

Preferred skills

GCP/GKE, AWS/EKS, ML multi-node/multi-cluster workloads, shipping to customer-hosted Kubernetes

Technologies

OpenTelemetry, Kubernetes, Terraform, Helm, Ray

Responsibilities

Evolve observability stack for logs, metrics, traces, and alerting; Ensure dataplanes report health, heartbeat, logs, metrics, and usage; Detect drift between desired and actual state; Monitor deployment and runtime agent health; Define SLOs and lead incident response/postmortems; Reduce telemetry costs

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.