Data Engineer
Core
Designing, building, and optimizing scalable data pipelines, warehouses, and lakes to support business intelligence and machine learning initiatives.
Role type
Data Engineer
Builds
Scalable data infrastructure (pipelines, warehouses, lakes) for analytics and ML
Domain
Cloud data engineering and big data processing
Deliverable
production ML models | infrastructure
Required skills
SQL, Python, Scala, cloud data solutions (AWS Redshift, Google BigQuery, Azure Synapse, Snowflake), ETL pipeline development (Apache Airflow, dbt, Talend, Fivetran), data modeling, schema design, database optimization, big data frameworks (Apache Spark, Hadoop, Kafka, Flink), containerization (Docker, Kubernetes), CI/CD workflows, data security and governance
Preferred skills
Experience with real-time and batch data processing, orchestration tools, debugging large-scale data challenges
Technologies
AWS Redshift, Google BigQuery, Azure Synapse, Snowflake, Apache Airflow, dbt, Talend, Fivetran, Apache Spark, Hadoop, Kafka, Flink, Docker, Kubernetes
Responsibilities
Design and maintain scalable data pipelines and ETL workflows; Develop and optimize data warehouses and data lakes; Implement real-time and batch data processing solutions; Work with structured and unstructured data for modeling and storage; Ensure data reliability, consistency, and scalability; Collaborate with analysts and scientists for efficient data access; Automate data ingestion, transformation, and validation; Monitor and optimize query performance; Implement security, compliance, and governance standards; Stay updated with emerging data engineering trends
Seniority
Mid-level (4+ years experience)