Staff Software Engineer - Observability
Core
Design and implement metrics, logging, tracing, and alerting infrastructure to enable fast debugging and high reliability for large-scale, performance-critical distributed AI inference systems.
Role type
Staff Software Engineer (Observability & Distributed Systems)
Builds
Internal observability platforms, telemetry pipelines, libraries, and tooling for AI inference services
Domain
AI/ML Infrastructure, Distributed Systems, Observability
Deliverable
production ML models
Required skills
Backend/systems software engineering, Go/C++/Rust/Java/Python, Distributed systems, Networking fundamentals, Concurrency and performance tradeoffs, Metrics/logs/distributed tracing, Production monitoring and alerting, OpenTelemetry, Prometheus, Grafana, Datadog/Elastic/Jaeger/Tempo, High-signal alerts design, Scalable telemetry pipelines, SLIs/SLOs design
Preferred skills
High-performance computing, AI/ML systems, Inference platforms, Hardware-aware observability, SRE or platform engineering background, Large-scale production incident debugging, Internal developer platforms
Technologies
OpenTelemetry, Prometheus, Grafana, Datadog, Elastic, Jaeger, Tempo
Responsibilities
Design and implement observability instrumentation across services and platforms; Build and maintain telemetry pipelines for metrics, logs, and traces at scale; Develop internal observability platforms, libraries, and tooling; Define and operationalize SLIs, SLOs, and alerting strategies; Partner with engineers to make systems debuggable by design; Reduce MTTR by enabling fast root-cause analysis during incidents
Seniority
Staff, hands-on IC with strategic impact