CareerPlanGet AI match score →

Data Engineer (Python)

🌐 Remote💼 Full-time🗓 2026-07-30

Core

Building and managing large enterprise data and analytics platforms, including scalable data lakes, ingestion pipelines, and hyper-scale processing clusters.

Role type

Senior Data Engineer (Python)

Builds

Scalable Smart Data Lakes, Data Ingestion Platforms, Machine Learning and NLP based Analytics Platforms, Hyper-Scale Processing Clusters, Data Mining and Search Engines

Domain

Big Data, Data Engineering, Cloud Infrastructure

Deliverable

production ML models | infrastructure

Required skills

Python, PySpark, Pandas, SQL, NoSQL, Stream-processing (Kafka, Spark-Streaming), Event-driven architectures, RESTful API development

Preferred skills

Unit testing (pytest), Docker, Kubernetes

Technologies

PySpark, Pandas, Postgres, MongoDB, Elasticsearch, Kafka, Spark-Streaming, Docker, Kubernetes, pytest

Responsibilities

Create and manage end-to-end data solutions and optimal data processing pipelines for large volume, big data sets; Develop efficient pre-processing and data manipulation tasks; Implement and manage scalable data lakes and hyper-scale processing clusters; Work with data science and infrastructure teams to implement practical machine learning solutions and pipelines in production.

Seniority

Mid-Senior, hands-on IC

Rewrite
# Data Engineer (Python) ## Location Gurgaon/Hyderabad/Remote ## About the role We are expanding our Data Engineering Team and hiring passionate professionals with extensive knowledge and experience in building and managing large enterprise data and analytics platforms. We are looking for creative individuals with strong programming skills, who can understand complex business and architectural problems and develop solutions. The individual will work closely with the rest of our data engineering and data science team in implementing and managing Scalable Smart Data Lakes, Data Ingestion Platforms, Machine Learning and NLP based Analytics Platforms, Hyper-Scale Processing Clusters, Data Mining and Search Engines. ## Responsibilities - Create and manage end-to-end Data Solutions, Optimal Data Processing Pipelines and Architecture dealing with large volume, big data sets of varied data types. - Work with data science and infrastructure team members to implement practical machine learning solutions and pipelines in production. ## Requirements - 3+ years of industry experience in creating and managing end-to-end Data Solutions, Optimal Data Processing Pipelines and Architecture dealing with large volume, big data sets of varied data types. - Proficiency in Python. Knowledge of design patterns, oops concepts and strong design skills in python. - Strong knowledge of working with PySpark Dataframes, Pandas Dataframes for writing efficient pre-processing and other data manipulation tasks. - Experience with creating Restful web services and API platforms. - Experience with SQL and NoSQL databases. E.g. Postgres/MongoDB/ Elasticsearch etc. - Experience with stream-processing systems: Spark-Streaming, Kafka etc and working experience with event-driven architectures. ## Good to have - Experience with testing libraries such as pytest for writing unit-tests for the developed code. - Knowledge and experience of Dockers and Kubernetes would be good to have.
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗