CareerPlanGet AI match score →
💼 Full-time🗓 2026-06-25

Core

Design and maintain declarative extraction specifications and build autonomous, self-healing data extraction pipelines using AI agents and classical scraping backends.

Role type

Senior AI Data Engineer (Extraction Engineering)

Builds

Reusable specification libraries, component-based extraction platforms, and autonomous data extraction agents.

Domain

Data Engineering / Web Scraping / AI Automation

Deliverable

production ML models | product features

Required skills

Python, Specification-Driven Extraction, LangChain, LangGraph, LlamaIndex, AutoGen, Scrapy-LLM, Playwright + MCP, Autonomous Agent Design, Classical Scraping Fundamentals, Data Validation & Storage, API Integration, HTTP/DOM/XPath/CSS

Preferred skills

Open-source contributions to scraping/AI-automation, Data privacy engineering (GDPR/CCPA), DevOps (Docker, CI/CD)

Responsibilities

Design declarative extraction specifications using Pydantic/JSON schemas; Implement pipelines translating specs into executable plans; Deploy self-healing spiders using Model Context Protocol (MCP); Orchestrate multi-step browsing workflows with agentic frameworks; Build component-based extraction platforms with monitoring and alerting; Partner with data scientists to refine specs for complex domains.

Seniority

Senior, hands-on IC

Rewrite
## About the Role Design and maintain declarative extraction specifications—using Pydantic models, JSON schemas, or domain-specific languages—that describe exactly which fields to capture, their types, and validation rules. Implement pipelines that translate these specifications into executable extraction plans, leveraging both classical (Scrapy, Playwright) and AI-augmented (LLM-based semantic parsing) backends. Build reusable specification libraries for recurring data types (product prices, tariff codes, regulatory texts) to accelerate onboarding of new sources. ## Responsibilities ### Autonomous & Self-Healing Systems * Deploy self-healing spiders that automatically detect website layout changes and repair themselves using Model Context Protocol (MCP) servers (e.g., Scrapy MCP Server, Playwright MCP). * Integrate semantic extraction (Scrapy-LLM, custom LLM pipelines) to eliminate selector brittleness—spiders rely on field descriptions, not fragile XPaths. * Hands-on experience building AI agents and orchestration systems. * Orchestrate complex, multi-step browsing workflows with agentic frameworks (BMAD/TEA, AutoGPT-like agents) that reason about page state, adapt to anti-bot measures, and correct their own behaviour in real time. ### Platform Thinking & Reusability * Move beyond one-off scrapers: build a component-based extraction platform where selectors, login handlers, and pagination logic are shared, versioned, and tested. * Implement monitoring, alerting, and automatic rollback for failed extraction runs. * Champion ethical crawling by design—rate limiting, robots.txt respect, and compliance with GDPR/CCPA are built into the specification layer, not retrofitted. ### Collaboration & Continuous Innovation * Partner with data scientists and domain experts to refine extraction specifications for complex, unstructured domains (e.g., legal texts, tariff classifications). * Evaluate and pilot emerging tools to push automation coverage beyond 90%. * Document and evangelise specification-driven best practices across the engineering organisation. ## Qualifications * Bachelor's degree in Computer Science * 3+ years of experience in web scraping or data extraction ## Required Skills * Proficiency with Python * Experience with specification-Driven Extraction * Experience with LangChain, LangGraph, LlamaIndex, AutoGen * Hands‑on use of Scrapy‑LLM, Scrapy MCP Server, or similar systems that decouple field definitions from page structure * Familiarity with frameworks that give LLMs browser control (Playwright + MCP, BMAD/TEA) to handle complex, non‑deterministic crawling tasks. * Design and implement autonomous data extraction agents that can make decisions about source selection, retry logic, and parsing strategies * Classical Scraping Fundamentals * Data Validation & Storage – Ability to define validation rules within specifications and land clean data into SQL/NoSQL databases or data lake * Basic API integration and authentication flows. * HTTP, DOM, XPath, CSS. ## Nice to Haves * Contributions to open-source scraping or AI-automation projects. * Contributions to open-source scraping or AI-automation projects. * Familiarity with data privacy engineering (GDPR, CCPA) baked into specification design. * DevOps light – Docker, CI/CD for testing extraction specifications. ## Mindset & Approach (Non-Negotiable) * Strong belief that the future of scraping is declarative, not imperative. * Candidate rather write a schema that says "extract the price" than debug an XPath when a website redesigns. * Looking to shift from "code that scrapes" to "systems that understand extraction"
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗