Principal GenAI Data Engineer
Core
Architecting enterprise-grade Generative AI data ingestion, knowledge preparation, and platform architectures to enable scalable, production-ready GenAI applications.
Role type
Principal GenAI Data Engineer
Builds
Enterprise data platforms for ingesting, processing, governing, and serving structured and unstructured data for AI/LLM workloads.
Domain
Cybersecurity / Generative AI / Enterprise Data Engineering
Deliverable
production ML models
Required skills
Enterprise data architecture, Unstructured data pipelines, GenAI platform engineering, Python, Distributed/scalable data pipelines, Vector databases, Graph databases, Metadata/knowledge storage systems, Clustering algorithms, Entity recognition, RAG workflows
Preferred skills
Real-time distributed vector search infrastructure, Multi-modal knowledge graph pipelines, LLMOps / GenAIOps frameworks, Agent Frameworks
Technologies
LangSmith, Arize Phoenix, Weights & Biases, MLflow, LangGraph, CrewAI, Google ADK
Responsibilities
Architect enterprise-scale GenAI data platforms for ingestion, transformation, enrichment, and serving; Design scalable pipelines for enterprise knowledge ingestion from diverse data sources; Define architecture for metadata extraction, chunking, enrichment, embeddings generation, and knowledge preparation workflows; Design AI-ready data models and storage strategies for vector, graph, and hybrid knowledge systems; Architect scalable unstructured data processing pipelines for text, images, PDFs, tables, and multimodal content.
Seniority
Principal, hands-on IC & architecture