Data Engineer
Core
Building and operating production AI data infrastructure including ingestion pipelines, storage design, and retrieval layers for RAG systems.
Role type
Senior Data Engineer (AI Infrastructure)
Builds
Ingestion/transformation pipelines, data warehouses/lakehouses, and retrieval infrastructure for AI models.
Domain
AI Engineering / Data Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Python, SQL, Spark, Kafka, dbt, Airflow, Snowflake, BigQuery, AWS, Azure, Docker, CI/CD, Terraform, Vector databases (pgvector, Pinecone)
Preferred skills
Flink, Dagster, Prefect, Jupyter, HIPAA compliance, Open-source contributions
Technologies
Spark, Kafka, dbt, Airflow, Snowflake, BigQuery, Redshift, Databricks, pgvector, Pinecone, Qdrant, Azure AI Search, Azure, AWS, GitHub Actions, Terraform
Responsibilities
Design idempotent batch and streaming ingestion pipelines; Architect warehouse and lakehouse storage solutions; Build chunking, embedding, and indexing pipelines for RAG systems; Implement data quality testing and lineage; Manage sensitive data security and access controls; Operate containerized deployments with cost optimization.
Seniority
Senior, hands-on IC