Data Developer
Core
Build data pipelines, RAG systems, and scalable data solutions for AI and ML-powered trading tools.
Role type
Data Developer (AI/ML Infrastructure)
Builds
Data pipelines for RAG systems, centralized data lakes, and vector storage solutions for AI researchers.
Domain
Financial Trading / AI & Machine Learning
Deliverable
production ML models
Required skills
RAG architectures, vector databases, Python, DAG-based orchestration, embedding models, distributed data processing, LLM inference optimization, containerization, data modeling, ETL/ELT patterns
Preferred skills
Semantic search systems, GPU-accelerated inference optimization, data versioning, caching strategies
Technologies
Milvus, ChromaDB, Pinecone, Weaviate, Qdrant, Airflow, Dagster, Prefect, Apache Spark, Ray, Dask, Docker
Responsibilities
Design and build data pipelines for RAG systems including document ingestion and vector storage; Build ingestion pipelines for structured and unstructured data into a centralized data lake; Develop data processing workflows to prepare datasets for fine-tuning and inference; Build monitoring and evaluation frameworks for retrieval quality and system performance; Collaborate with ML engineers to optimize data formats for GPU-accelerated inference; Implement caching strategies and data versioning systems; Deploy and manage vector databases and embedding services.
Seniority
Mid-level, hands-on IC