Senior Software Engineer, Data Engineering
Core
Design, develop, and maintain data ingestion, transformation, and orchestration pipelines; build and optimize data Lakehouse architectures; deploy and maintain data APIs and ML models in production.
Role type
Senior IC data engineer (cloud, data lakehouse, MLOps)
Builds
Data Lakehouse architectures, ETL/ELT pipelines, data APIs, and ML models for life sciences clients
Domain
Life Sciences / Data Engineering / Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis
Required skills
Python, SQL, C#/Java/Scala, Azure Data Lake, Databricks, Azure Data Factory, Apache Airflow, dbt, Kafka, Event Hubs, Service Bus, TensorFlow, Pytorch, Power BI, Apache Superset, Tableau, Azure, Docker, Kubernetes, GitHub Actions, Azure DevOps
Preferred skills
knowledge graph, semantic search, LLM-based data retrieval (RAG), data mesh, data fabric, MLflow, Delta Live Tables, Unity Catalog
Responsibilities
Design, develop, and maintain data ingestion, transformation, and orchestration pipelines (batch and real-time); Build and optimize data Lakehouse architectures using Azure Synapse, Delta Lake, or similar frameworks; Integrate and manage structured and unstructured data sources (SQL/NoSQL, files, documents, IoT streams); Develop and operationalize ETL/ELT pipelines using Azure Data Factory, Databricks, or Apache Spark; Collaborate with Data Scientists to prepare and serve ML-ready datasets for model training and inference; Implement data quality, lineage, and governance frameworks across pipelines and storage layers; Deploy and maintain data APIs and ML models in production using Azure ML, Kubernetes, and CI/CD pipelines
Seniority
Senior, hands-on IC