Senior Data Scientist - Protein Data Pipelines
Core
Building scalable, reliable data pipelines and MLOps frameworks to transform protein property data into ML-amenable assets for predictive modeling of sequence, structure, and function.
Role type
Senior IC data scientist (computational biology/MLOps)
Builds
Production-ready data pipelines, inference/deployment frameworks, and reusable validation systems for protein discovery models
Domain
Computational biology / Bioinformatics / Machine Learning
Deliverable
production ML models
Required skills
Python, SQL, Databricks, MLOps, model lifecycle management, data quality control, reproducibility practices, cross-functional collaboration
Preferred skills
MLflow, wet-lab collaboration, protein sequence/structure datasets, deployment workflows
Technologies
Python, SQL, Databricks, MLflow
Responsibilities
Design and maintain scalable data pipelines for protein property data; Develop deployment and inference strategies for machine learning models; Establish data quality, validation, and monitoring practices; Mediate collaborations between ML developers and wet-lab scientists; Document data lineage and reproducibility practices
Seniority
Senior, hands-on IC