Senior Data Engineer
Core
Owns the end-to-end data layer for Generative AI products, building pipelines that feed retrieval systems, agents, and analytics.
Role type
Senior Data Engineer (Generative AI)
Builds
Data pipelines, retrieval layers, and curated datasets for AI/LLM systems
Domain
Generative AI, Data Engineering, Cloud Infrastructure
Deliverable
production ML models
Required skills
SQL, Python, PySpark, AWS data stack (S3, Glue, Athena, Redshift), layered data architecture (lakehouse/medallion), ELT tools (Airbyte/Fivetran), event-driven pipelines (SQS/SNS/Kinesis), semantic layers (dbt/Cube), vector stores (pgvector/Pinecone), open table formats (Iceberg/Delta Lake)
Preferred skills
Infrastructure as Code, DevOps collaboration
Technologies
AWS, PySpark, Airbyte, Fivetran, Kinesis, Amazon MSK, dbt, Cube, pgvector, Pinecone, Apache Iceberg, Delta Lake, Hudi, MWAA, Step Functions, Dagster, Prefect
Responsibilities
Build and run batch and streaming pipelines from source to warehouse; Build data layer for retrieval including chunking, embedding, and vector indexing; Model curated datasets and metrics for AI consumers; Implement quality checks, validation, and monitoring; Apply access control and PII handling; Expose data services with platform/DevOps; Optimize storage, compute, and query costs; Review code and document standards
Seniority
Senior, hands-on IC