CareerPlanGet AI match score →

Lead Data Engineer

💼 Full-time🗓 2026-07-27

Core

Designing, building, and maintaining scalable data infrastructure (pipelines, warehouses, systems) to power AI-native tools for scientific discovery.

Role type

Lead Data Engineer

Builds

Data pipelines, data warehouses, and systems for AI products

Domain

AI/ML infrastructure, scientific discovery

Deliverable

production ML models

Required skills

ETL/ELT design, cloud data warehousing, data modeling, performance optimization, data quality frameworks

Preferred skills

AI/ML infrastructure, vector databases, streaming data, biotechnology domain knowledge

Technologies

Cloud data warehouses, ETL/ELT tools

Responsibilities

Design and maintain scalable ETL/ELT pipelines; Own database performance (schema, indexing, query optimization); Build and improve data warehousing infrastructure; Establish data quality standards and monitoring; Partner with research and product teams; Lead architectural decisions on new data systems

Seniority

Lead, hands-on IC with architectural ownership

Rewrite
## About the role We're looking for a Lead Data Engineer to own the data infrastructure that powers GXL's AI products (eg. Paperclip). You'll be responsible for designing, building, and maintaining the pipelines, warehouses, and systems that our research and products depend on. ## Responsibilities - Design and maintain scalable ETL/ELT pipelines across structured and unstructured data sources - Own database performance: schema design, indexing strategies, query optimization, and capacity planning - Build and improve data warehousing infrastructure - Establish and enforce data quality standards, validation frameworks, and monitoring - Partner closely with research and product teams to understand data requirements and ship reliably - Lead architectural decisions on new data systems ## Requirements - 2+ years of data engineering experience, with a track record of owning complex pipelines end-to-end - Deep experience with ETL/ELT design patterns and tools - Experience with at least one major cloud data warehouse - Solid understanding of data modeling, normalization, and warehousing best practices - Experience diagnosing and resolving performance bottlenecks at scale - Comfort operating in ambiguous, fast-moving environments ## Nice to have - Experience in AI/ML infrastructure or working alongside model training pipelines - Familiarity with vector databases or embedding pipelines - Experience with streaming data - General understanding of biotechnology industry ## About the company Generative Expert Labs (GXL) is building the next generation of AI-native tools for scientific discovery. We're a fast-moving team that ships products that agents and humans love, and we're looking for engineers who want to disrupt the existing agentic ecosystem.
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗