Data Engineer - Generative AI & Vector Systems
Core
Design and build scalable data pipelines to prepare, transform, and index enterprise data for Generative AI applications, semantic search, and RAG systems.
Role type
Senior Data Engineer (Generative AI & Vector Systems)
Builds
Scalable data ingestion pipelines, vector database architectures, and production-grade ETL/ELT workflows for LLM and Voice AI solutions.
Domain
Generative AI, Vector Search, Cloud Data Engineering
Deliverable
production ML models
Required skills
Python (Advanced), SQL (Advanced), ETL/ELT, Data Transformation, Metadata Management, Apache Airflow, AWS services (S3, Glue, EMR, Lambda, Athena), Vector Databases (Pinecone, Milvus, Qdrant, Chroma, Weaviate, FAISS), Starburst/Trino, LLMs, Text Embeddings, Semantic Search, RAG
Preferred skills
Prompt Engineering, REST APIs, JSON, Data Connectors
Technologies
Python, SQL, Apache Airflow, Pinecone, Milvus, Qdrant, Chroma, Weaviate, FAISS, AWS S3, AWS Glue, Amazon EMR, Lambda, Athena, Starburst, Trino, Presto
Responsibilities
Design and develop scalable data ingestion pipelines for AI/ML applications; Build automated pipelines to clean, transform, chunk, enrich, and load enterprise data into Vector Databases; Design and manage Vector Database architectures and optimize indexing/retrieval performance; Develop production-grade ETL/ELT pipelines with batch and near real-time ingestion; Write complex SQL queries across distributed data sources and design efficient data models for AI workloads; Work closely with AI engineers to support LLM-based applications and build RAG pipelines.
Seniority
Mid-to-Senior, hands-on IC