AI Data Engineer
Core
Build and operate production data pipelines and infrastructure to support AI applications and knowledge layers at HHMI, transforming institutional content into governed, AI-ready data.
Role type
Senior hands-on AI Data Engineer
Builds
AI-facing data pipelines, medallion architecture tables, workflow orchestration frameworks, and retrieval-supporting infrastructure (embeddings, vector stores)
Domain
Biomedical research + AI data engineering
Deliverable
production ML models
Required skills
Production data engineering, Databricks (Delta Lake, DLT, Workflows, Unity Catalog), Python/PySpark, SQL, Workflow orchestration, Data quality & observability, Governance patterns, AWS foundations (via careerplan.io/jobs/R-4645-ai-data-engineer-at-hhmi)
Preferred skills
Knowledge graphs, Modern data stack, LLM-based extraction frameworks, Cross-platform data engineering
Technologies
Databricks, Delta Lake, Delta Live Tables, Unity Catalog, PySpark, Terraform, AWS (S3, IAM, KMS), Git, CI/CD
Responsibilities
Build ingestion, transformation, and serving pipelines for AI-ready data; Implement medallion architecture patterns; Own workflow orchestration and alerting; Implement governance and access control; Build and maintain embedding/vector store infrastructure; Define source-system contracts; Operate data quality checks and drift detection
Seniority
Senior, hands-on IC
