Principal Data Engineer
Core
Design and operate the end-to-end data platform transforming raw vendor files into finished datasets for B2B buyers, including transformation stacks, matching systems, and internal curation applications.
Role type
Principal Data Engineer (Staff-level IC)
Builds
Production-grade data platform, entity matching systems, and data quality layers for Fortune 500 clients
Domain
B2B data intelligence, AI-driven datasets, and GTM analytics
Deliverable
production ML models | product features
Required skills
Data platform architecture, Spark, Databricks, SQL, dimensional modelling, OLTP (MySQL/Postgres), Airflow, AWS (EC2, S3, EMR), entity resolution, fuzzy matching, Python/Scala/Java, CI/CD, Infrastructure as Code
Preferred skills
B2B firmographic/technographic data, streaming/CDC, warehouse-to-lakehouse migration, AI/ML integration, Spark MLlib, Databricks ML, data governance (GDPR/CCPA)
Technologies
Spark, Databricks, Airflow, AWS (EC2, S3, EMR), MySQL, Postgres, Python, Scala, Java, Docker, Kubernetes, Terraform
Responsibilities
Own data platform architecture and build-vs-buy decisions; build quality layers for data freshness and completeness; improve entity matching and attribution; automate release cycles; optimize compute costs across Spark and serving tiers
Seniority
Principal, hands-on IC