CareerPlanSign in
Hyderabad, Telangana💼 Full-time🗓 2026-06-06 → 2026-08-09

Core

Designing and maintaining scalable data pipelines for processing large volumes of structured and unstructured data to support NLP and LLM applications.

Role type

Data Engineer (NLP/LLM pipelines)

Builds

Scalable data processing pipelines, document ingestion workflows, and vector database systems for AI/ML teams.

Domain

Artificial Intelligence / Natural Language Processing / Large Language Models

Deliverable

production ML models

Required skills

Python, Apache Spark, Spark SQL, Hugging Face Transformers, PDF parsing, OCR, HTML parsing, text extraction, document chunking, vector databases, Git, CI/CD, ETL workflows

Preferred skills

Text normalization, sentence segmentation, deduplication, data masking, data classification, embedding generation, RAG pipelines, performance optimization, CUDA optimization, PyTorch tuning

Technologies

Python, Apache Spark, Spark SQL, Hugging Face, PDF libraries, OCR tools, Vector Databases, Git, CI/CD tools

Responsibilities

Design and develop scalable data pipelines for structured and unstructured data; Build document ingestion and processing workflows for PDFs, scanned documents, and HTML; Implement OCR, PDF parsing, and text extraction pipelines; Develop document chunking and preprocessing frameworks for NLP and LLM applications; Create and optimize data transformation workflows using Python and Apache Spark; Develop and manage Vector Database pipelines for embedding storage and retrieval; Collaborate with AI/ML engineers to prepare datasets for model training and inference; Monitor, troubleshoot, and improve production data processing systems.

Seniority

Mid-level, hands-on IC

Sourced via adzuna · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.