CareerPlanSign in

Staff Software Engineer, Observability & Profiling

London, UK💼 Full-time💰 $325,000–$325,000🗓 2026-09-02 → 2026-09-25

Core

Design and build scalable telemetry ingest and storage pipelines, observability solutions, and instrumentation libraries for Anthropic's massive GPU, TPU, and Trainium clusters to ensure operational reliability.

Role type

Staff Software Engineer (Infrastructure/Observability)

Builds

High-throughput telemetry pipelines, fleet-wide continuous profiling, eBPF-based tracing, and AI-assisted diagnostic tooling

Domain

Cloud Infrastructure, Observability, High-Performance Computing (GPU/TPU/Trainium)

Deliverable

production ML models | infrastructure

Required skills

Large-scale observability infrastructure design, end-to-end telemetry signal management, high-throughput pipeline architecture, kernel/network stack debugging, eBPF development, continuous profiling at fleet scale, accelerator workload instrumentation, OpenTelemetry expertise

Preferred skills

eBPF-based observability in production, fleet-scale continuous profiling management, kernel/syscall-level debugging, accelerator profiling, high-cardinality metrics systems, AI/LLM application to operational workflows

Technologies

eBPF, OpenTelemetry, GPU/TPU/Trainium clusters, kernel stack, network stack

Responsibilities

Design scalable telemetry pipelines for metrics, logs, traces, and error data; Build observability solutions for deep system visibility; Own and evolve core observability platforms; Build instrumentation libraries and SDKs; Reduce MTTR via cross-signal correlation and unified query interfaces; Drive fleet-wide efficiency through continuous profiling insights

Seniority

Staff, hands-on IC with strategic impact

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.