Data Scientist II
Core
Build computational infrastructure and reproducible pipelines to support multi-omics analysis and accelerate cancer research at the National Cancer Institute.
Role type
Senior IC data scientist (bioinformatics & multi-omics)
Builds
Reproducible data pipelines, interactive dashboards, and curated datasets for systems-level biological questions
Domain
Biomedical informatics, translational research, and multi-omics data science
Deliverable
production ML models | dashboards & analysis | infrastructure
Required skills
Python, R, bulk RNA-seq, single-cell RNA-seq, spatial transcriptomics, statistical modeling, dimensionality reduction, data integration, cloud data platforms, workflow automation
Preferred skills
scverse ecosystem, Digital Spatial Profiling, HPC environments, HL7/FHIR standards, training researchers
Technologies
Python, R, Shiny, Streamlit, AWS, Snowflake, Databricks, Nextflow, WDL, Snakemake, Singularity, Terra, Galaxy
Responsibilities
Design and maintain reproducible pipelines for genomic, transcriptomic, and clinical datasets; support multi-omics analysis workflows including QC, clustering, and differential expression; build researcher-facing dashboards and APIs; enable data ingestion, curation, and lifecycle management; apply statistical and ML methods to biomedical data; collaborate with scientists to translate needs into technical solutions
Seniority
Mid-Senior, hands-on IC