Principal Engineer - Data Ingestion & AI Pipeline
Core
Designing and leading scalable data ingestion pipelines to prepare enterprise data for RAG and AI workloads, ensuring data is clean, traceable, secure, and optimized for LLM consumption.
Role type
Principal Engineer - Data Ingestion & AI Pipeline
Builds
Scalable ingestion pipelines, transformation patterns, and orchestration frameworks for structured, semi-structured, and unstructured data.
Domain
Financial Services / Enterprise AI & Data Engineering
Deliverable
production ML models
Required skills
Large-scale enterprise system design, LLM/RAG/vector search/AI platform capabilities, production-grade data pipeline design, ETL/ELT/streaming/batch processing, cloud data platforms (Azure Data Factory, Synapse, Databricks, Spark, Kafka, Airflow), document processing/OCR/parsing, data governance/lineage/security patterns, architecture leadership across multiple teams.
Preferred skills
Azure AI Document Intelligence/Microsoft Graph/SharePoint/Purview, large-scale data lakehouse architectures, semantic chunking/entity extraction/knowledge graphs, sensitive data detection/PII handling/DLP.
Technologies
Azure Data Factory, Synapse, Databricks, Fabric, Spark, Kafka, Event Hubs, Airflow, dbt, Azure AI Document Intelligence, Microsoft Graph, SharePoint, Purview.
Responsibilities
Develop engineering approach for program/portfolio solutions; lead planning and design of complex features spanning multiple teams; define technology tool stacks; lead technical oversight including design reviews; establish ingestion standards for lineage, freshness, and access controls; implement pipeline orchestration, monitoring, and error handling; partner with RAG and context engineers to optimize data for embedding and reasoning; evaluate and implement OCR and document intelligence capabilities; ensure compliance with security and regulatory requirements; mentor senior engineers and influence technical direction.
Seniority
Principal, hands-on IC with strategy & mentorship