CareerPlanGet AI match score →

Site Reliability Engineering Tech Lead

Hybrid - Palo Alto🌐 Remote💼 Full-time🗓 2026-03-20 → 2026-07-31

Core

Lead technical initiatives for DataHub Cloud and enterprise deployment solutions to ensure reliability, scalability, and operational excellence for an AI & Data Context Platform.

Role type

Senior SRE Tech Lead

Builds

Robust management control planes, automated deployment pipelines, monitoring/alerting systems, and self-service tools for enterprise-grade DataHub deployments.

Domain

AI & Data Infrastructure / Cloud Platforms

Deliverable

production ML models | infrastructure

Required skills

Cloud platforms (AWS, GCP, Azure), Containerization (Docker, Kubernetes), Infrastructure as Code (Terraform, CloudFormation, Pulumi), Python/Java, Monitoring/Observability (Prometheus, Grafana, Datadog), CI/CD pipelines, Networking, Security, Database operations.

Preferred skills

Multi-tenant SaaS platforms, Customer-facing deployment tools, Data infrastructure/metadata management, Service mesh/microservices, Enterprise client interaction, Data governance.

Responsibilities

Design scalable infrastructure solutions, Lead multi-cloud deployment strategies, Architect monitoring/observability systems, Drive IaC and automation best practices, Partner on advanced deployment capabilities, Establish and maintain SLAs/SLOs, Lead incident response and post-mortems, Implement chaos engineering, Mentor SRE engineers, Improve on-call practices and knowledge sharing.

Seniority

Senior, hands-on IC with leadership responsibilities

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗