CareerPlanGet AI match score →

Senior Lead Data Engineer

💼 Full-time🗓 2026-07-25

Core

Building scalable real-time and batch ETL/ELT pipelines for manufacturing analytics using high-frequency industrial data.

Role type

Senior Lead Data Engineer

Builds

Cloud-based data architectures (data lakes, lakehouses, warehouses) and optimized data processing workflows.

Domain

Manufacturing analytics, Industrial IoT, Cloud Data Engineering

Deliverable

production ML models

Required skills

PySpark, Azure Databricks, Python, Apache Spark, SQL, NoSQL (MongoDB, Cassandra), Docker, Kubernetes, CI/CD, Medallion Architecture

Preferred skills

MLOps, DevOps, model lifecycle management, time series databases (Influx DB)

Technologies

Azure Databricks, PySpark, Apache Spark, Azure, AWS, GCP, SQL Server, PostgreSQL, Influx DB, MongoDB, Cassandra, Docker, Kubernetes

Responsibilities

Build scalable real-time and batch processing workflows; Design and maintain cloud-based data architectures; Deploy and optimize data solutions on cloud platforms; Develop and optimize ETL/ELT pipelines from IoT, MES, SCADA, LIMS, and ERP systems; Automate data workflows using CI/CD; Monitor, troubleshoot, and enhance data pipelines.

Seniority

Senior, hands-on IC with team leadership

Rewrite
## About the role We are looking for a Senior Data Engineer with strong expertise in Azure Databricks, PySpark, and distributed computing to develop and optimize scalable ETL pipelines for manufacturing analytics. The role involves working with high-frequency industrial data to enable real-time and batch data processing. ## Key Responsibilities - Build scalable real-time and batch processing workflows using Azure Databricks, PySpark, and Apache Spark. - Perform data pre-processing, including cleaning, transformation, deduplication, normalization, encoding, and scaling to ensure high-quality input for downstream analytics. - Design and maintain cloud-based data architectures, including data lakes, lakehouses, and warehouses, following Medallion Architecture. - Deploy and optimize data solutions on Azure (preferred), AWS, or GCP with a focus on performance, security, and scalability. - Develop and optimize ETL/ELT pipelines for structured and unstructured data from IoT, MES, SCADA, LIMS, and ERP systems. - Automate data workflows using CI/CD and DevOps best practices, ensuring security and compliance with industry standards. - Monitor, troubleshoot, and enhance data pipelines for high availability and reliability. - Utilize Docker and Kubernetes for scalable data processing. - Collaborate with automation team, data scientists and engineers to provide clean, structured data for AI/ML models. ## Desired Skills and Qualifications - Bachelor's or Master's degree in Computer Science, Information Technology, or a related field from Tier 1 institutes. (IIT, NIT, IIIT, DTU etc.) - 5+ years of experience in core data engineering, with a strong focus on cloud platforms such as Azure (preferred), AWS, or GCP. - Proficiency in PySpark, Azure Databricks, Python and Apache Spark, etc. - 2 years of team handling experience. - Expertise in relational databases (e.g., SQL Server, PostgreSQL), time series databases (e.g. Influx DB), and NoSQL databases (e.g., MongoDB, Cassandra). - Experience in containerization (Docker, Kubernetes). - Strong analytical and problem-solving skills with attention to detail. - Good to have MLOps, DevOps including model lifecycle management. - Excellent communication and collaboration skills, with a proven ability to work effectively as a team player. - Comfortable working in a dynamic, fast-paced startup environment, adapting quickly to changing priorities and responsibilities.
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗