Data Engineer
Core
Designing, optimizing, and automating large-scale distributed data pipelines to ensure data quality and enable sharing across the organization.
Role type
Senior Data Engineer
Builds
Scalable batch data pipelines and data sharing mechanisms
Domain
Data Engineering / Big Data
Deliverable
production ML models | product features | infrastructure
Required skills
PySpark, SQL optimization, Python, Airflow, Livy, Flink, Apache Iceberg, DataHub, Grafana
Preferred skills
AI-powered developer assistants, data quality frameworks
Technologies
PySpark, SQL, Airflow, Livy, EMR, Flink, Iceberg, Snowflake, DataHub, Grafana, SFTP
Responsibilities
Optimize existing PySpark pipelines for performance and scalability; Tune SQL queries for execution time and resource utilization; Design and develop new batch data pipelines; Automate pipeline generation and metadata management; Implement data quality and validation frameworks; Integrate and share data via DataHub and Snowflake Data Share; Monitor pipelines using Grafana and telemetry tools; Participate in design reviews and operational support.