Senior Site Reliability Engineer I
Core
Build and evolve Axon's next-generation observability platform (tracing, logging, metrics) to enable engineering teams to operate services with confidence.
Role type
Senior Site Reliability Engineer (Observability Infrastructure)
Builds
Distributed tracing infrastructure, log aggregation platform, metrics infrastructure, and self-service observability tooling.
Domain
Cloud-native observability, SRE, Infrastructure Engineering
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Linux systems fundamentals, Kubernetes, Infrastructure as Code (Terraform, CDK), Golang/Python/Java, LGTM stack (Loki, Grafana, Tempo/Jaeger, Cortex), agentic AI tooling, distributed tracing adoption
Preferred skills
OpenTelemetry expertise, GitOps workflows, debugging complex multi-service distributed systems, experience with 24/7 high-volume systems
Technologies
OpenTelemetry, Jaeger, Loki, Alloy, Cortex, Prometheus, Grafana, Terraform, CDK, ArgoCD, Helm, Kubernetes
Responsibilities
Own and evolve distributed tracing infrastructure; Build and operate log aggregation platform; Maintain metrics infrastructure; Write internal tooling and automation for self-service observability; Manage infrastructure as code and participate in on-call rotation; Partner with engineering teams to define instrumentation standards and drive SLO adoption
Seniority
Senior, hands-on IC