Data Engineer, AI Support
Core
Build reliable data workflows and prepare datasets for AI/ML research projects, ensuring data quality, reproducibility, and responsible handling.
Role type
Data Engineer (AI Support)
Builds
Scalable data pipelines, curated datasets for LLM/VLM training, and integrated data solutions for research teams.
Domain
Artificial Intelligence / Machine Learning / Data Engineering
Deliverable
production ML models
Required skills
Python (vectorized libraries), SQL/NoSQL, data modeling, ETL pipeline development, dataset versioning, data quality validation, Linux/Bash scripting, CI/CD automation, distributed processing concepts.
Preferred skills
Cloud platforms (Azure/AWS/GCP), orchestration tools (Airflow/Prefect/Dagster/Spark), data annotation platforms (Label Studio/Prodigy), experiment tracking (Weights & Biases/Langfuse).
Technologies
Python, SQL, NoSQL, Linux, Bash, GitHub Actions, GitLab CI, Terraform, Weights & Biases, Langfuse, Azure, AWS, Google Cloud, Airflow, Prefect, Dagster, Spark, Label Studio, Prodigy.
Responsibilities
Prepare and organize data for research and machine learning projects; build data workflows moving raw data to usable datasets; identify quality issues and structure data for AI applications; automate data preparation tasks; ensure datasets are consistent, traceable, and reproducible; integrate data across research tools and AI platforms; protect sensitive information and follow data handling practices.
Seniority
Mid-level, hands-on IC