Staff Site Reliability Engineer
Core
Architect and maintain resilient, scalable infrastructure for a global agentic software creation platform serving millions of developers.
Role type
Staff Site Reliability Engineer (IC)
Builds
Production-grade observability solutions, automation frameworks, and self-healing systems for a cloud-native platform.
Domain
Cloud-native infrastructure, distributed systems, and developer tools.
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Python or Go, Kubernetes, distributed systems architecture, observability (metrics/logging/tracing), incident management, Infrastructure as Code (Terraform/Pulumi), capacity planning, debugging complex distributed systems.
Preferred skills
Google Cloud Platform (GCP), Prometheus, Grafana, Datadog, OpenTelemetry, rapid-growth startup environment experience.
Technologies
Kubernetes, Terraform, Pulumi, Python, Go, GCP, Prometheus, Grafana, Datadog, OpenTelemetry, Docker.
Responsibilities
Design and lead implementation of comprehensive monitoring, logging, and tracing solutions; define and drive Service Level Objectives (SLOs) and Service Level Indicators (SLIs); lead incident response and conduct blameless post-mortems; architect and build automation to eliminate toil; optimize performance on large-scale Kubernetes deployments; debug and harden distributed systems; review feature and system designs for reliability and scalability; mentor and educate the broader engineering team.
Seniority
Staff, hands-on IC with strategic guidance and mentorship responsibilities.