Associate Scientist, Data II
Core
Design, build, and maintain semantic data infrastructure and knowledge graphs to connect information across domains, translating complex source data into well-modeled, interlinked graph structures.
Role type
Associate Scientist, Data II (Knowledge Graph Engineering)
Builds
Knowledge graphs, ingestion pipelines, and query systems for data and analytics solutions.
Domain
Healthcare / Life Sciences (Immunology, Oncology, Neuroscience) + Knowledge Graph Engineering
Deliverable
production ML models | infrastructure
Required skills
Graph data modeling, pipeline development, query engineering, data quality and validation, performance tuning, collaboration
Preferred skills
Entity resolution, record linkage, knowledge graph construction, vector databases, embeddings, RAG patterns
Technologies
Neo4j, Amazon Neptune, Tiger Graph, RDF triple stores, Cypher, SPARQL, SQL, Python, Apache Spark, SAS, R
Responsibilities
Design and refine labeled-property and/or RDF graph models; Build, test, and maintain ingestion pipelines; Write, optimize, and document Cypher queries; Implement constraints, validation rules, and automated tests; Monitor graph performance and tune indexes; Partner with cross-functional teams to gather requirements and document designs.
Seniority
Associate (0-2 years experience with Bachelor's or entry-level with Master's)