Senior Data Engineer (AI/ML)
Core
Build scalable data and AI infrastructure for next-generation intelligent products, combining modern data engineering with Generative AI to support training, inference, and retrieval workloads.
Role type
Senior IC Data Engineer (AI/ML)
Builds
Production-grade data platforms, RAG systems, and AI applications using LLMs and vector search.
Domain
Generative AI, Large Language Models (LLMs), Data Engineering, Distributed Systems
Deliverable
production ML models | product features | infrastructure
Required skills
Python, Scala/Java, SQL, Databricks, Apache Spark, Delta Lake, Snowflake, Airflow, LLMs, RAG, embeddings, vector search, streaming architectures, ETL/ELT, data governance, observability
Preferred skills
LangGraph, LangChain, LlamaIndex, vector databases (Qdrant, Pinecone, Weaviate), Kafka, MLflow, Unity Catalog, model-serving platforms
Technologies
Databricks, Apache Spark, Delta Lake, Snowflake, Airflow, Kafka, Spark Structured Streaming, LangGraph, LangChain, LlamaIndex, Qdrant, Pinecone, Weaviate, Databricks Vector Search, MLflow, Unity Catalog
Responsibilities
Design and build AI/LLM data pipelines for training, inference, and retrieval; Build production-grade RAG systems covering ingestion, chunking, embedding, indexing, and reranking; Develop AI applications using LLMs, structured outputs, and agentic workflows; Build and optimize semantic search and vector retrieval systems; Develop frameworks for LLM evaluation, monitoring, and cost optimization; Design scalable batch and streaming pipelines; Establish data quality, governance, lineage, and security practices; Partner with ML teams to transition prototypes to production; Monitor production AI systems for performance, latency, and health.
Seniority
Senior, hands-on IC