AI Data Engineer
Core
Design and maintain declarative extraction specifications and build autonomous, self-healing data extraction pipelines using AI agents and classical scraping backends.
Role type
Senior AI Data Engineer (Extraction Engineering)
Builds
Reusable specification libraries, autonomous self-healing spiders, and component-based extraction platforms.
Domain
Data Engineering / AI Automation / Web Scraping
Deliverable
production ML models
Required skills
Python, Specification-Driven Extraction, LangChain, LangGraph, LlamaIndex, AutoGen, Scrapy-LLM, Playwright + MCP, Autonomous Agent Design, Classical Scraping Fundamentals, Data Validation & Storage, API Integration, HTTP/DOM/XPath/CSS
Preferred skills
Open-source contributions to scraping/AI-automation, Data privacy engineering (GDPR/CCPA), DevOps (Docker, CI/CD)
Responsibilities
Design declarative extraction specifications using Pydantic/JSON schemas; Implement pipelines translating specs to executable plans; Deploy self-healing spiders using Model Context Protocol; Orchestrate multi-step browsing workflows with agentic frameworks; Build component-based extraction platforms with monitoring and rollback; Partner with data scientists to refine specs for unstructured domains.
Seniority
Senior, hands-on IC