Site Reliability Engineer III
Core
Build and maintain the infrastructure platform enabling developers to ship reliably at scale for millions of outdoor enthusiasts.
Role type
Senior Site Reliability Engineer (Infrastructure)
Builds
Highly available systems, deployment automation, and observability for onX's outdoor mapping products.
Domain
Cloud Infrastructure & DevOps
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, Terraform/OpenTofu, Cloud SQL, Bigtable, Google Composer (Airflow), Google Cloud Storage, BigQuery, Pub/Sub, Cloud Run, Google Cloud Monitoring, Prometheus, OpenTelemetry, Checkly, Rootly, SQL/NoSQL datastores, networking, incident response, architectural decision-making
Preferred skills
Google Cloud Platform expertise, troubleshooting high-throughput/low-latency services, IAM/security management, GIS mapping systems, Airflow/ETL systems, Claude Code
Technologies
Terraform, CockroachDB, GCP, GKE, Cloud SQL, Bigtable, Composer, Cloud Storage, BigQuery, Pub/Sub, Cloud Run, Prometheus, OpenTelemetry, Checkly, Rootly
Responsibilities
Deploy, monitor, and maintain highly available systems; maintain and extend Terraform codebase; analyze systems for performance, availability, and cost optimization; automate manual systems; develop integrations with monitoring/alerting tools; drive incident response best practices; collaborate on architectural decisions
Seniority
Senior, hands-on IC