DevOps
Core
Drive reliability, scalability, and operational excellence for DataHub Cloud and enterprise deployment solutions, ensuring AI and data system reliability.
Role type
Senior DevOps/Site Reliability Engineering (SRE) Engineer
Builds
Production systems for AI-native data context platform, including deployment automation, monitoring, and self-healing capabilities
Domain
Enterprise SaaS, AI infrastructure, Data governance
Deliverable
production ML models | infrastructure
Required skills
Cloud platforms (AWS, GCP, Azure), Containerization (Docker, Kubernetes), Infrastructure as Code (Terraform, CloudFormation), Python/Java programming, CI/CD pipelines, Monitoring/Observability (Prometheus, Grafana, Datadog)
Preferred skills
Multi-tenant SaaS platform operations, Customer-facing deployment tool development, Data infrastructure and metadata management
Responsibilities
Partner with product/engineering to influence advanced deployment capabilities; Establish and maintain SLAs/SLOs; Lead incident response and post-mortems; Optimize system performance, capacity planning, and cost efficiency; Improve on-call practices and runbooks
Seniority
Senior, hands-on IC
