Data Engineer
Core
Design and develop batch and streaming data processing pipelines on Google Cloud Platform to support Data Science and Analytics teams.
Role type
Data Engineer (Big Data & GCP)
Builds
Production data pipelines and optimized datasets on BigQuery
Domain
Big Data Engineering, Google Cloud Platform
Deliverable
production ML models | product features
Required skills
Apache Beam, Apache Spark, Google Cloud Dataproc, BigQuery (advanced SQL), Python, SQL (advanced), Git/CI-CD
Preferred skills
Google Cloud Dataflow, Pipeline orchestrators (Apache Airflow/Cloud Composer), Data modeling (star/snowflake schema, data vault), Google Cloud Professional Data Engineer certification
Technologies
Apache Beam, Apache Spark, Google Cloud Dataproc, BigQuery, BigQuery Studio, Python, Git, CI-CD, Google Cloud Dataflow, Apache Airflow, Cloud Composer
Responsibilities
Design and develop batch and streaming data processing pipelines using Apache Beam and Apache Spark on Google Cloud Dataproc; Model, optimize, and query large datasets on BigQuery with cost and performance focus; Integrate heterogeneous data sources (relational databases, APIs, event streams) into GCP pipelines; Monitor data quality and implement tests and alerts on production pipelines; Collaborate with Data Science teams to ensure data availability and reliability; Contribute to team data engineering standards including naming conventions and data lineage
Seniority
Mid-level (2–4 years experience)