Data Engineer / 1
Core
Design, develop, and deploy Python-based ETL/ELT pipelines to migrate data from on-premises MS SQL Server into Databricks for advanced AI and ML projects.
Role type
Data Engineer
Builds
Data pipelines and migrated datasets for AI/ML initiatives
Domain
Cloud Data Engineering / AI & Machine Learning
Deliverable
production ML models
Required skills
Python, PySpark, Pandas, Databricks, Spark ecosystem, ETL/ELT concepts, data modeling, pipeline orchestration, Microsoft SQL Server, Parquet ingestion, Delta Lake, structured streaming
Preferred skills
AI/ML data preparation workflows, data governance and compliance, orchestration tools (Databricks Workflows, Airflow), Databricks environment setup
Technologies
Databricks, Delta Lake, MS SQL Server, Parquet, Python, PySpark, Pandas
Responsibilities
Design and deploy ETL/ELT pipelines; implement data validation and quality assurance checks; tune pipeline performance and manage resource usage; collaborate with AI engineers and data scientists on data access patterns; maintain technical documentation for pipelines.
Seniority
Mid-level, hands-on IC