Lead Software Engineer, Data Engineering
Core
Architect, design, and maintain data ingestion, transformation, and orchestration pipelines (batch and real-time) for life sciences companies.
Role type
Lead Software Engineer, Data Engineering
Builds
Data Lakehouse architectures, ETL/ELT pipelines, data APIs, and ML models
Domain
Life Sciences / Data Engineering
Deliverable
production ML models | data infrastructure
Required skills
Python, SQL, C#/Java/Scala, Azure Data Lake, Databricks, Azure Data Factory, Apache Airflow, dbt, Kafka, Event Hubs, Service Bus, TensorFlow, Pytorch, Power BI, Apache Superset, Tableau, Azure, Docker, Kubernetes, GitHub Actions, Azure DevOps
Preferred skills
knowledge graph, semantic search, LLM-based data retrieval (RAG), data mesh, data fabric, MLflow, Delta Live Tables, Data Bricks Unity Catalog
Responsibilities
Architect and develop data ingestion, transformation, and orchestration pipelines; Build and optimize data Lakehouse architectures; Integrate and manage structured and unstructured data sources; Develop and operationalize ETL/ELT pipelines; Collaborate with Data Scientists to prepare ML-ready datasets; Implement data quality, lineage, and governance frameworks; Deploy and maintain data APIs and ML models in production; Ensure scalability, performance, and observability of data workflows; Mentor junior team members
Seniority
Lead, hands-on IC with mentorship