Data & Knowledge Engineer
Core
Design and build ingestion, transformation, and serving pipelines for structured and unstructured data to ground agentic workflows and improve their reliability.
Role type
Senior IC data & knowledge engineer (AI/ML infrastructure)
Builds
Retrieval indexes, metadata models, semantic layers, knowledge graphs, and data products
Domain
Enterprise data engineering, AI/ML infrastructure, information retrieval
Deliverable
production ML models | infrastructure
Required skills
SQL, Python, data pipelines, APIs, data modeling, vector search, embeddings, metadata management, document processing, retrieval evaluation, enterprise security, hybrid data integration
Preferred skills
Spark, Microsoft Fabric, Azure Data Factory, Databricks, Snowflake, OCR, chunking, enrichment, lineage, incremental indexing, graph technologies (Neo4j, RDF), ontologies, entity resolution, GraphRAG patterns, reranking, query transformation
Technologies
Spark, Microsoft Fabric, Azure Data Factory, Databricks, Snowflake, Azure AI Search, PostgreSQL (pgvector), Elasticsearch, Pinecone, Weaviate, Milvus, Neo4j, RDF
Responsibilities
Design and build ingestion, transformation, and serving pipelines; Create retrieval indexes, metadata models, semantic layers, and knowledge graphs; Implement chunking, enrichment, lineage, quality, and access-control patterns; Optimize retrieval quality, freshness, latency, and cost; Integrate cloud and on-premises data sources for hybrid solutions; Support evaluation datasets, monitoring data, and traceability requirements
Seniority
Senior, hands-on IC