Microbiologist IV (Genomic Data Engineer)
Core
Build and maintain distributed genomic data pipelines for pathogen genomics, public health surveillance, and outbreak detection.
Role type
Senior Genomic Data Engineer (Public Health)
Builds
Scalable data engineering workflows for epidemiological and laboratory datasets
Domain
Public Health / Genomics / Big Data
Deliverable
production ML models | infrastructure
Required skills
Hadoop ecosystem (HDFS, Spark, Hive, Impala), ETL development, genomic data ingestion, data validation & harmonization, Git version control, Python/Scala/Rust/Bash scripting, data governance
Preferred skills
Pathogen genomics experience, Spark-based analytics, bioinformatics workflows, cloud data platforms, data lineage documentation
Technologies
Hadoop, Spark, Hive, Impala, Git, Python, Scala, Rust, Bash, NCBI GenBank, SRA
Responsibilities
Develop and optimize distributed data pipelines; Manage large-scale ETL workflows for genomic/epidemiological datasets; Ingest and harmonize data from external repositories; Ensure compliance with data governance for sensitive health data; Collaborate with bioinformaticians and epidemiologists; Prepare technical documentation and reports
Seniority
Senior, hands-on IC