CareerPlanGet AI match score →

Senior Ai Data Engineer

🌐 Remote💼 Full-time💰 $1–$1🗓 2026-07-24

Core

Building the data layer and pipelines that enable AI agents to ingest, process, and reason over vast unstructured market datasets for investment decisions.

Role type

Senior AI Data Engineer (Agentic Systems)

Builds

Data pipelines, knowledge graphs, and retrieval systems optimized for AI agent consumption

Domain

Asset Management / Financial Data Engineering / Agentic AI

Deliverable

production ML models | infrastructure

Required skills

Data pipeline design, knowledge graph modeling, vector embeddings, data quality monitoring, schema drift detection, graph database querying, agentic context management

Preferred skills

Financial data literacy, rapid prototyping, system design for automated consumers

Technologies

Graph databases (Neo4j, Kuzu, NetworkX), vector databases, unstructured data processing

Responsibilities

Design and implement pipelines ingesting earnings transcripts, SEC filings, and alternative data; model entity relationships using graph structures; build monitoring for data quality and observability; optimize data structures for agent retrieval within context windows

Seniority

Senior, hands-on IC

Rewrite
## About the role As an AI Data Engineer, you'll build the data layer that allows our agents to make market-beating investment decisions. You will design and implement pipelines that ingest, process, and structure large volumes of financial data including earnings transcripts, SEC filings, analyst estimates, financial research, pricing data, and alternative sources like social media, news feeds, and community forums. Your work product is optimized for consumption by AI agents, not humans. You will model relationships between data points (companies, competitors, suppliers, macro indicators, etc.) using graph structures that allow agents to traverse and reason over complex market dynamics. Candidates must have a deep understanding of how agents consume and reason across data and design systems accordingly. You will own considerable portions of our data stack and must be comfortable working in a rapidly iterative environment. We're looking for engineers who are immersed in agents and capable of designing customized systems that support and empower human investors. Critical to our process is your ability to make data available to our agentic reasoning systems. ## What You Have - Passion for building data systems for agents, not people. You understand that agents retrieve and consume data differently than humans do. They depend entirely on what is surfaced to them, cannot fill gaps with intuition, and fail silently when retrieval is poor. You develop proof of concepts, make quick decisions, and deliver on accelerated deadlines. - Mastery of data structures optimized for agent retrieval. You understand the tradeoffs between graphs, vector embeddings, tabular databases, indexes, and other data-related concepts. - You treat retrieval quality as a first-class engineering concern, not an afterthought. - You have experience building and querying knowledge graphs to represent entity relationships. You think explicitly about how each structure performs when an agent is the query originator, and design accordingly. - Agentic context and memory: Humans can reference a document they read last week. Agents cannot. Data must be structured and surfaced in ways that fit within context windows, with knowledge extracted and stored explicitly for future retrieval. - Data quality and observability: At scale, silent data failures are more dangerous than loud ones. You build monitoring into your pipelines such as tracking coverage gaps, schema drift, embedding degradation, and graph consistency so that agents reason on data you can vouch for. - Financial Literacy: You don't need to be a finance expert, but you must have an interest in learning. - 2+ years of professional data engineering experience or a strong portfolio of data-intensive agentic personal projects. - Candidates must be currently residing in the US with valid work authorization or OPT status. H-1B sponsorship is evaluated on a case-by-case basis. ## Short Answer Questions - Complex Pipeline: Describe a pipeline you built that processed large volumes of data for downstream consumption by an automated system. What retrieval or search mechanism did you use, what tradeoffs did you make to optimize for accuracy, and how did your design choices account for the fact that the consumer was a system, not a person. - Experience with knowledge graphs (Optional): Have you worked with knowledge graphs or graph databases (e.g., Neo4j, Kuzu, NetworkX)? If so, describe the problem you were modeling and why a graph structure was the right choice. ## About the company About Alka Intelligence We are an asset management startup building a proprietary reasoning engine to automate fundamental analysis. We are building Alka Arbiter, a system capable of ingesting vast, unstructured market datasets and transforming them into clear investment mandates. Alka Arbiter will be used to manage our own funds. Unlike standard AI wrappers, our engine is designed to show its work, creating a fully auditable record of why a decision was made. We are looking for builders to help us bridge the gap between stochastic AI models and the rigorous reliability required to manage fiduciary capital. Locations: San Francisco, Dallas, and Remote Compensation: $200,000–$250,000 + 0.5–1.0% equity + Pro-Rata Management Fees, commensurate with seniority. (Lower bound adjusted on platform for seniority)
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗