AI Data Engineer
Core
Design and build high-throughput, low-latency data pipelines, feature stores, and ML infrastructure to power production AI and LLM applications for global financial institutions.
Role type
Senior Associate AI/GenAI Data Engineer
Builds
Production-grade data systems, feature stores, vector databases, and ML pipelines for LSEG's product suite
Domain
Financial markets infrastructure and data
Deliverable
production ML models
Required skills
Python, Node.js, SQL (ClickHouse, PostgreSQL, Snowflake), NoSQL (Elasticsearch, Redis), Apache Kafka, AWS, CI/CD (GitHub Actions, Jenkins), Docker, Kubernetes, Terraform, Helm, LangChain, LlamaIndex, Scikit-learn, XGBoost, PyTorch, MLflow
Preferred skills
MCP (Model Context Protocol), RAG pipelines, embedding models, fine-tuning LLMs, prompt engineering, LLM evaluation, async programming, FastAPI, PySpark, Kafka Streams, financial market event processing, cost optimization, multi-region architecture
Technologies
OpenAI, Anthropic, LangChain, LlamaIndex, Scikit-learn, XGBoost, PyTorch, MLflow, ClickHouse, PostgreSQL, Snowflake, Elasticsearch, Redis, Apache Kafka, GitHub Actions, Jenkins, Docker, Kubernetes, Terraform, Helm, AWS (S3, RDS, Redshift, Lambda, ECS/EKS, SageMaker)
Responsibilities
Design and operate high-throughput, low-latency data pipelines; Build and maintain feature stores, vector databases, and data layers for production ML models; Architect SQL and NoSQL data systems for analytical and operational workloads; Implement Redis-based caching and pub/sub systems; Build and deploy machine learning pipelines from preprocessing to serving; Integrate and fine-tune large language models; Design and own CI/CD pipelines for data and ML systems
Seniority
Senior Associate, hands-on IC