Senior Python Data Engineer (OCR & Document Processing)- remote
Core
Design, build, and optimize scalable data ingestion and document processing solutions to transform unstructured insurance data into structured, AI-ready information for downstream AI and retrieval systems.
Role type
Senior Python Data Engineer (OCR & Document Processing)
Builds
Cloud-native data pipelines, OCR/document extraction workflows, and vector database schemas for RAG solutions.
Domain
Insurance / Document Intelligence / Cloud Data Engineering
Deliverable
production ML models | infrastructure
Required skills
Python, SQL, AWS (S3, Step Functions, CloudWatch), OCR/document extraction, vector databases, RAG concepts, CI/CD, automated testing
Preferred skills
Azure cloud services, Databricks, experience in regulated industries (Insurance, Banking)
Technologies
AWS, Python, SQL, SharePoint, PDF, Word, Excel, PowerPoint, Email, Git
Responsibilities
Design scalable data ingestion pipelines for unstructured documents; Integrate and optimize OCR technologies; Build automated workflows for parsing and metadata enrichment; Develop connectors for enterprise repositories; Design vector database schemas for RAG; Implement monitoring and quality-control mechanisms; Optimize workflows for scalability and low-latency.
Seniority
Senior, hands-on IC