Principal Knowledge & Data Architect
Core
Design and operate the knowledge management layer for AI, turning institutional information into structured, retrievable knowledge for retrieval-augmented generation (RAG) pipelines and knowledge graphs.
Role type
Principal Knowledge & Data Architect
Builds
RAG pipelines, knowledge graphs, and entity resolution systems for HHMI's administrative and operational functions.
Domain
Biomedical research institutions / Generative AI / Knowledge Engineering
Deliverable
production ML models
Required skills
Production RAG systems, Knowledge graph design and operation, Entity resolution and deduplication, Information extraction (LLM/NLP/Rule-based), Data modeling and classification, Python, SQL, Vector databases, Graph databases
Preferred skills
Hybrid retrieval and GraphRAG patterns, Master data management, Semantic web technologies (RDFS, OWL, SKOS), AI agent tool design
Technologies
Python, SQL, Postgres pgvector, Pinecone, Weaviate, Qdrant, Neo4j, AWS Neptune, JanusGraph, TigerGraph, Stardog, Cypher, SPARQL, Gremlin, Databricks, dbt
Responsibilities
Design knowledge representation strategies (chunking, embeddings, graphs), build and operate RAG pipelines, construct knowledge graphs for complex reasoning, extract structure from unstructured documents, solve entity resolution and canonicalization, govern knowledge classification and lineage, partner with data integrations and AI platform teams
Seniority
Principal, hands-on IC with strategy & mentorship