Data Engineer – AWS + Hadoop
Core
Design and implement scalable ETL/ELT pipelines, data lakes, and distributed data platforms on AWS and Hadoop to support analytics and machine learning.
Role type
Senior Data Engineer (AWS + Hadoop)
Builds
Production data pipelines, data lakes, and curated datasets for analytics and ML teams.
Domain
Cloud data engineering, big data processing, distributed systems.
Deliverable
production ML models | product features | dashboards & analysis
Required skills
AWS data services (S3, Glue, EMR, Redshift), Hadoop ecosystem (HDFS, Hive, Spark, Kafka), PySpark/Scala, SQL, data modeling, CI/CD, Docker, Git, IAM/security governance
Preferred skills
Lake Formation, cost optimization, advanced Spark tuning, curated data APIs
Technologies
AWS (S3, Glue, EMR, Athena, Lambda, Redshift, IAM, CloudWatch), Hadoop (HDFS, Hive, Spark, Kafka, Oozie, Airflow), Docker, Jenkins, GitHub Actions
Responsibilities
Design and implement batch and streaming ETL/ELT pipelines; Build ingestion frameworks using Kafka/Kinesis and Spark; Develop and optimize AWS-based data lakes and warehouses; Manage Hadoop ecosystem tools and job orchestration; Implement data quality, governance, and access controls; Monitor pipelines and improve cost, performance, and reliability
Seniority
Senior, hands-on IC