CareerPlanGet AI match score →

Data Infrastructure Engineer Ai Agents

💼 Full-time🗓 2026-07-25

Core

Building data infrastructure and AI agents to transform unstructured financial documents into structured datasets for institutional buyers.

Role type

Senior IC Data Infrastructure Engineer (AI Agents)

Builds

Data pipelines, AI agents, and structured datasets for financial services

Domain

Financial services, AI agents, data infrastructure

Deliverable

production ML models | product features | dashboards & analysis

Required skills

Python, Postgres, SQLite, async programming, queues and task managers, web scraping, LLMs in production, data pipeline architecture, ETL/ELT patterns, data validation and quality monitoring

Preferred skills

JavaScript/TypeScript, GCP/AWS, C/C++, Rust, Go, ML pipeline tooling (Airflow, Dagster, Prefect), feature engineering, model fine-tuning

Technologies

Claude Code, Codex, Gas Town, OpenClaw

Responsibilities

Own the entire data wrangling function (ingestion, cleanup, storage, transformation, distribution), build and maintain AI agents, ensure pipeline reliability and data quality, make architecture decisions

Seniority

Mid-Senior, hands-on IC with CTO trajectory potential

Rewrite
## About the role Role currently for humans only. Full-time, in-office — Old Street, London. We build structured datasets from unstructured financial documents. AI agents do the reading. We build the agents. The team is five people. The output looks like a department of fifty. We're hiring an engineer to own data wrangling. The entire function. Ingestion, cleanup, storage, transformation, distribution. You'll work with existing AI agents and build new ones. The goal: reinvent how datasets are built, maintained, and delivered in financial services when the workforce is mostly artificial. ## The work You will build the infrastructure that turns 100,000+ unstructured PDFs into constantly updating datasets that institutional buyers rely on for their decision-making. You'll own it. Architecture decisions, pipeline reliability, data quality, delivery format. If something breaks at 2am, it's yours. If something ships to a client faster than anyone expected, also yours. ## Fundamentals You're genuinely excited about this. Not just interested. Energised. You love building things that matter. You're internally driven by the AI agent space and want to be part of something that's actually changing how work gets done. Maybe you've started experimenting with AI agents on your own. Maybe you've already got OpenClaw up and running as your personal assistant. You thrive in a tight, small team where everyone ships. You've got energy and ambition, but you're not naive about what it takes. You care deeply about data quality. Not in a checkbox way. In a "this inconsistency is going to bother me until I fix it" way. You understand that the datasets we build inform financial decisions that move millions of dollars. A missing data point isn't trivial to you. An anomaly isn't something to gloss over. You have integrity about data. You're persistent, thorough, and you've done this before. Show us the projects where you cared about getting it right. You've built data pipelines or ML systems. Or you're obsessed with learning how. You have a data scientist's mindset: you think about problems end-to-end, from raw input to reliable output. Maybe you've built production pipelines. Maybe you've trained models. Maybe you've done serious data analysis work. Or maybe you're earlier in your career but you pick things up fast and you're genuinely hungry to go deep. Either way, you think like someone who understands what good data infrastructure looks like. ## Specifics You've shipped in a small team. Startup, or a small function inside something bigger. 2 to 30 people. You know what it means when there's no one to hand things off to. You use Claude Code or Codex daily — as your base infrastructure. If you've been deep in the agentic tooling ecosystem — building, testing what breaks, pushing boundaries — even better. Talk to us about Gas Town. You care about the outcome for the user. And not the elegance of the system. You care about shipping great data as a product and understand this process inherently. ## Experience 3–5 years is the likely range. Could be less if you're sharp. Could be more if you've kept moving. ## Skills that matter - Python - Postgres - SQLite - Async - Queues and task managers - Local servers - File I/O - Cron jobs - Web scraping - LLMs in production or serious side projects - Data pipeline architecture - ETL/ELT patterns - Data validation and quality monitoring ## Skills that might be useful - JavaScript/TypeScript - GCP/AWS - C/C++ - Rust - Go - ML pipeline tooling (Airflow, Dagster, Prefect) - Feature engineering - Model fine-tuning ## Skills we won't optimise for - Pandas wizardry - PyTorch - R - ML model training from scratch — although fine-tuning familiarity with local models is a plus ## Ownership You run with your strengths. When you hit something unfamiliar, you say so — then figure it out. You don't wait for tickets. You drive your own projects because that's how you work. Confidence without overcommitment. Know what you don't know. ## How we'll assess you You'll meet everyone and work with the team. We'll build a testing programme around you. We'll want to see how you think about data problems. Not just how you code them. If you built a 20k-star open source repo by yourself, we won't ask for a coding test — we'll test how you work with others. If you've managed a team of 3 elite engineers, we know you can collaborate — we'll want to see your ability to output features. We'll also want to understand: what's a data quality problem you caught that others missed? What's a pipeline you built that you're proud of? What gets you excited about this space? ## Trajectory This is a foundational role. If you're 3–5 years in and ambitious, this is a CTO path. If you're earlier, it's the fastest education in data infrastructure and AI agents you'll find. If you're later, it's a chance to own something end-to-end without the politics.
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗