Principal Site Reliability Engineer
Core
Establishing and driving Site Reliability Engineering practices from the ground up to ensure reliability, automation, and operational excellence for a platform serving millions of users.
Role type
Principal Site Reliability Engineer (Infrastructure & Observability)
Builds
Resilient, scalable cloud-native infrastructure and comprehensive observability platforms on Google Cloud Platform
Domain
Cloud Infrastructure / Site Reliability Engineering / Observability
Deliverable
production ML models | infrastructure
Required skills
Large-scale monitoring and observability design, distributed systems architecture, cloud-native patterns, infrastructure as code, deployment automation, Python/Go/Java programming, containerization and orchestration, CI/CD pipelines, database performance tuning, incident response management, cost optimization
Preferred skills
AIOps and machine learning for anomaly detection, establishing SRE programs from scratch, security and compliance frameworks, automated incident response and self-healing systems
Technologies
Google Cloud Platform, Kubernetes, Docker, Prometheus, Grafana, Datadog, Cloud Functions, GKE, SQL, NoSQL, Python, Go, Java, CI/CD
Responsibilities
Define SLIs, SLOs, error budgets, and reliability metrics; develop incident response protocols and post-mortem procedures; design disaster recovery and business continuity plans; lead architectural decisions for monitoring solutions; build custom SRE tools for automated monitoring and remediation; construct observability platforms; implement cloud-native monitoring systems; design auto-scaling and self-healing systems; optimize cloud costs; establish security and compliance frameworks; utilize AI/ML for predictive analytics; create actionable dashboards and reporting systems
Seniority
Principal, hands-on IC with leadership responsibilities