Data Engineering Intern
Core
Design, develop, and maintain data pipelines to ensure clean, reliable, and timely data for GenAI projects.
Role type
Data Engineering Intern
Builds
Data pipelines, ETL processes, data models, and visualizations for GenAI applications
Domain
Generative AI, Data Engineering
Deliverable
production ML models | infrastructure
Required skills
Python, SQL, data visualization, data cleaning, data validation, data transformation, problem solving
Preferred skills
AWS, PySpark, Airflow, data lakes, data warehouses, Git
Responsibilities
Design and develop data pipelines; implement and optimize ETL pipelines; integrate data from various sources into warehouses and lakes; support data management tasks; understand business objectives to develop data models; participate in client communications to gather requirements; identify data quality issues and improve codebase
Seniority
Intern