IND Staff Engineer, Reliability
Core
Staff Reliability Engineer ensuring the availability, scalability, and operational readiness of 200+ GCP services and data/ML workloads.
Role type
Staff IC Reliability Engineer (Cloud Infrastructure)
Builds
Production-grade GCP services, observability standards, automation tooling, and compliance frameworks for a multi-cloud environment.
Domain
Cloud Infrastructure / Data & Machine Learning Operations
Deliverable
production ML models | infrastructure
Required skills
GCP operations (Compute Engine, GKE, Cloud Run, BigQuery, Vertex AI), fault tolerance design, observability (SLI/SLO), incident response, automation (Cloud Functions, Cloud Run jobs), security governance (IAM, VPC, KMS), CI/CD (Terraform, Cloud Build), networking troubleshooting.
Preferred skills
Advanced GCP services (Dataflow, Pub/Sub, Cloud Composer), org-level policy creation, self-healing pattern design.
Technologies
Google Cloud Platform (GCP), BigQuery, Vertex AI, GKE, Cloud Run, Cloud Functions, Terraform, Cloud Build, Cloud Monitoring, Cloud Logging, Istio, Anthos.
Responsibilities
Operate and improve reliability for production GCP workloads; develop solutions for enterprise security and disaster recovery; drive automation for software-as-a-service efficiency; execute incident response and post-incident analysis; build observability standards and alerting strategies; partner with engineering teams on design reviews for operability.
Seniority
Staff, hands-on IC with strategic impact