Gen AI Data Engineer
Core
Designing and building distributed data systems and large-scale data warehouses to support scalable RAG solutions and generative AI workflows.
Role type
Senior Gen AI Data Engineer
Builds
Robust data pipelines for document ingestion, indexing, and retrieval; scalable RAG systems; data platforms processing petabytes of data.
Domain
Generative AI, Data Engineering, Knowledge Graphs, Vector Databases
Deliverable
production ML models | infrastructure
Required skills
Python, SQL, PySpark, Snowflake, Neo4j, Apache Airflow, AWS (S3, RDS, Lambda, SageMaker), GCP (Vertex AI, BigQuery), Linux, Docker, Kubernetes, Terraform, CI/CD (Jenkins, GitHub Actions), Vector Search (FAISS, Pinecone, Elasticsearch, Milvus), Graph Algorithms, Ontologies, Open Cypher, RDMBS, Unstructured Data Management (audio, video, image, text)
Preferred skills
Hadoop, Spark, Streamlit, Internal APIs, Basic ML concepts, Knowledge Graph creation/retrieval, Disaster Recovery planning, System monitoring and optimization
Technologies
Snowflake, Neo4j, Apache Airflow, AWS, GCP, Docker, Kubernetes, Terraform, Jenkins, GitHub Actions, FAISS, Pinecone, Elasticsearch, Milvus, Open Cypher
Responsibilities
Architecting data platforms for petabyte-scale processing; building pipelines for document ingestion and RAG support; curating data from diverse sources; integrating external databases and knowledge graphs into RAG systems; conducting experiments to evaluate and optimize RAG workflows; designing scalable infrastructure and implementing security best practices.
Seniority
Senior, hands-on IC