Site Reliability Engineering Tech Lead
Core
Lead technical initiatives for DataHub Cloud and enterprise deployment solutions to ensure reliability, scalability, and operational excellence for an AI & Data Context Platform.
Role type
Senior SRE Tech Lead
Builds
Robust management control planes, automated deployment pipelines, monitoring/alerting systems, and self-service tools for enterprise-grade DataHub deployments.
Domain
AI & Data Infrastructure / Cloud Platforms
Deliverable
production ML models | infrastructure
Required skills
Cloud platforms (AWS, GCP, Azure), Containerization (Docker, Kubernetes), Infrastructure as Code (Terraform, CloudFormation, Pulumi), Python/Java, Monitoring/Observability (Prometheus, Grafana, Datadog), CI/CD pipelines, Networking, Security, Database operations.
Preferred skills
Multi-tenant SaaS platforms, Customer-facing deployment tools, Data infrastructure/metadata management, Service mesh/microservices, Enterprise client interaction, Data governance.
Responsibilities
Design scalable infrastructure solutions, Lead multi-cloud deployment strategies, Architect monitoring/observability systems, Drive IaC and automation best practices, Partner on advanced deployment capabilities, Establish and maintain SLAs/SLOs, Lead incident response and post-mortems, Implement chaos engineering, Mentor SRE engineers, Improve on-call practices and knowledge sharing.
Seniority
Senior, hands-on IC with leadership responsibilities
