AI Data Engineer
Core
Designing, building, and maintaining robust data pipelines and infrastructure to power an internal AI platform (SchonAI) for investment professionals.
Role type
Senior AI Data Engineer
Builds
Scalable data pipelines, ETL/ELT processes, vector databases, and semantic search systems for LLM applications.
Domain
Financial services / AI Data Engineering
Deliverable
production ML models
Required skills
Python, SQL, distributed computing (Spark, Flink), cloud platforms (AWS/GCP), vector databases (Pinecone, Weaviate, Qdrant), embedding pipelines
Preferred skills
LLM/RAG pipeline experience, financial data sources, streaming technologies (Kafka, Kinesis), OLAP databases (BigQ, SingleStore, RedShift, ClickHouse), containerization (Docker, Kubernetes), MLOps, data quality frameworks (Great Expectations, Deequ)
Technologies
Prefect, Apache Airflow, Dagster, Kubernetes, AWS (S3), GCP, PostgreSQL, MySQL, MongoDB, DynamoDB, Elasticsearch, Kafka, Kinesis, Pub/Sub, Docker, RedShift, ClickHouse, Pinecone, Weaviate, Qdrant, Great Expectations, Deequ
Responsibilities
Design and build scalable data pipelines for structured and unstructured data; Develop ETL/ELT processes for diverse data sources; Implement real-time and batch data processing workflows; Build and maintain data infrastructure optimized for AI/ML workloads; Partner with AI engineers and data scientists to understand data requirements; Implement data access controls, encryption, and compliance measures.
Seniority
Senior, hands-on IC