Lead Data Engineer
Core
Architect and own the end-to-end data sourcing and ETL infrastructure powering a large-scale B2B data platform, transforming complex unstructured web data into accurate structured datasets.
Role type
Lead Data Engineer (AI-native data sourcing)
Builds
Production-grade extraction systems, agentic workflows, and LLM infrastructure for B2B data
Domain
B2B data, web extraction, AI/LLM infrastructure
Deliverable
production ML models | product features
Required skills
Python, SQL, Airflow, LLM infrastructure, agentic workflows, web extraction, schema design, cost optimization, observability, system architecture
Preferred skills
Snowflake, Redshift, AWS (Lambda, S3, ECS, Glue), vector databases, LLM evaluation tools, B2B data knowledge, GDPR/CCPA compliance
Technologies
Python, SQL, Airflow, Snowflake, Redshift, AWS, Claude Code, Cursor, Braintrust, Promptfoo, Inspect, Logfire, OpenTelemetry
Responsibilities
Architect and own the complete data sourcing and ETL layer; Develop and ship complex extractors handling anti-bot defenses and JavaScript-heavy sources; Build agentic extraction workflows with LLM escalation; Establish production-grade LLM infrastructure with versioning and evaluation; Design scalable orchestration using Airflow; Provide technical input to data platform and product teams; Contribute hands-on reference implementations for challenging extraction problems.
Seniority
Senior, hands-on IC with substantial autonomy