IND - Staff Engineer, Reliability
Core
Establishing and enforcing data reliability standards, monitoring data quality, and automating data pipeline operations to ensure data freshness, completeness, and accuracy.
Role type
Staff Engineer, Data Reliability & SRE
Builds
Reliable, observable, and self-healing data pipelines and data products
Domain
Insurance / Data Engineering & Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Data observability, data quality frameworks, pipeline automation, incident management, AIOps, prompt engineering, cloud engineering, distributed compute, data governance, DataOps
Preferred skills
Experience with Monte Carlo, Bigeye, Astro Observe, Datafold, AWS/GCP AI services
Technologies
Informatica, Python, PySpark, Amazon EMR, Hadoop, Snowflake, AWS, GCP
Responsibilities
Establish and enforce Data Service Level Objectives (SLOs) for data freshness and accuracy; Implement advanced data observability tools to monitor data journeys; Collaborate with Data Engineering to embed reliability patterns into pipelines; Automate data validation, reprocessing, and backfilling tasks; Lead response and resolution for data-related incidents; Develop and automate data-aware runbooks for pipeline failures and recovery scenarios