Data Engineer
Core
Designing and maintaining scalable data pipelines for processing large volumes of structured and unstructured data to support NLP and LLM applications.
Role type
Data Engineer (NLP/LLM pipelines)
Builds
Scalable data processing pipelines, document ingestion workflows, and vector database systems for AI/ML teams.
Domain
Artificial Intelligence / Natural Language Processing / Large Language Models
Deliverable
production ML models
Required skills
Python, Apache Spark, Spark SQL, Hugging Face Transformers, PDF parsing, OCR, HTML parsing, text extraction, document chunking, vector databases, Git, CI/CD, ETL workflows
Preferred skills
Text normalization, sentence segmentation, deduplication, data masking, data classification, embedding generation, RAG pipelines, performance optimization, CUDA optimization, PyTorch tuning
Technologies
Python, Apache Spark, Spark SQL, Hugging Face, PDF libraries, OCR tools, Vector Databases, Git, CI/CD tools
Responsibilities
Design and develop scalable data pipelines for structured and unstructured data; Build document ingestion and processing workflows for PDFs, scanned documents, and HTML; Implement OCR, PDF parsing, and text extraction pipelines; Develop document chunking and preprocessing frameworks for NLP and LLM applications; Create and optimize data transformation workflows using Python and Apache Spark; Develop and manage Vector Database pipelines for embedding storage and retrieval; Collaborate with AI/ML engineers to prepare datasets for model training and inference; Monitor, troubleshoot, and improve production data processing systems.
Seniority
Mid-level, hands-on IC