Lead Data Engineer (Cross Domain)
Core
Design and maintain robust data pipelines, drive deployment practices, and manage key initiatives in the Procurement domain.
Role type
Lead Data Engineer
Builds
End-to-end data products from source ingestion to Gold layer consumption
Domain
Procurement / Data Engineering
Deliverable
production ML models | product features | dashboards & analysis
Required skills
dbt, SQL, Python, PySpark, Databricks, Azure DevOps, GitHub, CI/CD, Infrastructure-as-Code, Data Vault 2.0, ETL, API ingestion
Preferred skills
Terraform, Bicep, scientific datasets (cheminformatics, bioinformatics, microbiome)
Technologies
dbt, Databricks, PySpark, Azure DevOps, GitHub Actions, Terraform, Bicep
Responsibilities
Design and implement end-to-end data pipelines across Bronze, Silver, and Gold layers; Develop and manage data ingestion processes from source systems; Build modular transformation frameworks using dbt and PySpark; Design, build, and maintain CI/CD and DevOps processes; Implement and manage Infrastructure-as-Code and deployment automation; Establish proactive monitoring, observability, and operational excellence practices; Drive data quality, governance, and compliance standards; Provide technical leadership and cross-functional collaboration
Seniority
Senior, hands-on IC with leadership responsibilities