CareerPlanGet AI match score →

Senior Data Engineer

💼 Full-time🗓 2026-07-26

Core

Design and build scalable modern data platforms, lakehouse architectures, and data infrastructure supporting analytics, AI, and machine learning workloads.

Role type

Senior IC data engineer

Builds

Production data platforms, ELT/ETL pipelines, and AI/ML data infrastructure

Domain

Cloud data engineering, distributed systems, AI/ML infrastructure

Deliverable

production ML models | product features | infrastructure

Required skills

Python (OOP), distributed data processing, ELT/ETL pipeline design, workflow orchestration, cloud platform management, infrastructure as code, data modeling, real-time streaming, vector databases, feature stores

Preferred skills

Lakehouse architectures (Iceberg, Delta Lake), CI/CD implementation, mentoring

Technologies

Python, Flask, FastAPI, dbt, Airflow, Dagster, Terraform, Docker, Kubernetes, Snowflake, Databricks, AWS Glue, GCP Composer, Dataflow, Dataproc, Apache Beam, Kafka, Flink, Apache Iceberg, Delta Lake, Medallion Architecture, vector databases

Responsibilities

Design and maintain scalable data platforms for analytics and AI; Build Lakehouse architectures; Develop data platform services and frameworks; Design, build, and optimize scalable ELT/ETL, batch, and real-time data pipelines; Develop and manage workflow orchestration and CI/CD pipelines; Build data pipelines for AI, ML, and RAG applications including embedding and feature store workflows

Seniority

Senior, hands-on IC

Rewrite
## About the Role We are seeking a highly skilled Senior Data Engineer with strong expertise in designing and building scalable modern data platforms. The ideal candidate will have extensive experience in data engineering, distributed data processing, real-time and batch pipelines, lakehouse architectures, and cloud-based analytics platforms. This role requires a self-driven engineer who can independently own technical initiatives, build reliable data infrastructure, collaborate with cross-functional teams, and contribute to the architecture and execution of enterprise-scale data platforms supporting analytics, AI, and machine learning workloads. ## Key Responsibilities ### 1. Data Platform Architecture & Engineering - Design and maintain scalable data platforms for analytics and AI workloads. - Build Lakehouse architectures using Apache Iceberg, Delta Lake, and Medallion Architecture. - Develop data platform services and frameworks using Python (Flask, FastAPI), applying Object-Oriented Programming (OOP) principles and strong problem-solving skills. - Implement data modeling, schema evolution, governance, and platform best practices. - Ensure platform scalability, reliability, security, and performance. ### 2. Data Engineering & Processing - Design, build, and optimize scalable ELT/ETL, batch, and real-time data pipelines using Python, dbt, Databricks, Snowflake, AWS Glue, GCP Composer, Dataflow, Dataproc, Apache Beam, Kafka, and Flink. - Develop distributed data processing solutions for large-scale datasets, ensuring high performance, reliability, and throughput. - Implement data quality, lineage, monitoring, validation, and governance processes across the data ecosystem. ### 3. Workflow Orchestration - Develop and manage workflows using Airflow and/or Dagster. - Implement CI/CD pipelines and DataOps best practices. - Manage infrastructure using Terraform, Docker, and Kubernetes. - Establish monitoring, alerting, and observability for production systems. ### 4. Collaboration & Ownership - Collaborate with Data, Software, AI/ML, and Product teams. - Own technical initiatives from design through deployment and maintenance. - Mentor team members and promote engineering best practices. ### 5. AI & Machine Learning Data Infrastructure - Build data pipelines for AI, ML, and RAG applications. - Develop embedding, vector database, and feature store workflows. - Support scalable data infrastructure for model training and inference. ## Qualifications - 4+ years of experience in Data Engineering. - B.E./B.Tech/B.S. Candidates' entries with significant prior experience in the fields above will be considered. - Strong proficiency in Python (Flask, FastAPI), Object-Oriented Programming (OOP), and problem-solving. - Proven experience building and managing large-scale ELT/ETL pipelines using dbt and Airflow/Dagster. - Hands-on experience with at least one cloud data platform (e.g., Snowflake, Databricks, or equivalent) and distributed data processing frameworks such as Apache Beam. - Good to have experience with Lakehouse architectures, including Apache Iceberg, Delta Lake, and Medallion Architecture. - Experience with real-time data streaming technologies such as Kafka and Apache Flink. - Experience building data infrastructure for AI/ML and Retrieval-Augmented Generation (RAG) applications. - Strong understanding of vector databases, embedding pipelines, and feature stores. - Experience with Terraform, Docker, Kubernetes, and CI/CD implementation. - Strong understanding of data quality, observability, monitoring, and operational best practices. - Excellent communication skills with the ability to work independently, mentor team members, and drive end-to-end project execution with minimal supervision. ## Work Location Ahmedabad/Pune ## How to Apply Contact us to apply. If you would like to apply for this role, send your resume to [email protected].
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗