Staff Software Engineer, Data Platform
Core
Building the central data platform frameworks, tooling, and paved paths that enable all teams at Harvey to work with data confidently and independently.
Role type
Staff Software Engineer, Data Platform
Builds
Ingestion layers, orchestration platforms, transformation/compute frameworks, self-serve tooling, stream processing infrastructure, and trust/governance layers.
Domain
Legal tech / Professional services / Cloud Data Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Cloud data warehouse architecture (Snowflake), CDC and streaming pipeline design, workflow orchestration at scale, data quality and observability, data governance and PII handling, framework and internal tooling development, schema evolution, multi-region data residency implementation.
Preferred skills
dbt, lakehouse architectures (Iceberg, Delta Lake), multi-tenant platform operations, data infrastructure for AI products.
Technologies
Snowflake, Kafka, Debezium, Flink, Spark Streaming, Fivetran, Airbyte, Temporal, Airflow, Dagster, Python, SQL, Azure, AWS, GCP, Kubernetes, Terraform, Pulumi.
Responsibilities
Own the data platform's architecture and technical direction; build and operate the ingestion layer across streaming, batch, and CDC; land data into Snowflake with defined freshness and cost characteristics; own the orchestration platform for scheduling and dependency management; build transformation and compute frameworks; design and operate stream processing infrastructure; build the trust layer for quality, lineage, and cataloging; implement patterns for PII and sensitive data handling; set the technical bar through design reviews and mentorship.
Seniority
Staff, hands-on IC with team leadership and technical direction responsibilities.