CareerPlanGet AI match score →

Data Scientist Applied Ml Engineer Arclogiq Software S Pvt Ltd

💼 Full-time💰 $10–$10🗓 2026-07-25

Core

Build and optimize DigiCatalog's probabilistic AI stack, focusing on selecting, orchestrating, and optimizing open-source AI models for OCR, multilingual NLP, computer vision, and multimodal pipelines.

Role type

Applied ML Engineer (Multimodal AI & OCR)

Builds

Scalable AI systems for document processing, image attribute extraction, and image enhancement

Domain

AI/ML, Computer Vision, NLP, Multimodal AI

Deliverable

production ML models

Required skills

Open-source AI model selection and orchestration, OCR pipeline development, Image attribute extraction, Image enhancement and generation, Multilingual AI workflows, Automated evaluation pipelines, Error analysis and calibration, ML system optimization for edge and cloud, Python, ML engineering, Inference optimization, Model benchmarking

Preferred skills

Edge AI deployments, Vector search and embeddings, GPU optimization, Distributed inference pipelines, Research mindset

Technologies

DocTR, PaddleOCR, TrOCR, Grounding DINO, OWL-ViT, CLIP, Small VLMs, Real-ESRGAN, SDXL-Turbo, IndicTrans2, NLLB, IndicWhisper

Responsibilities

Select and orchestrate open-source models based on use-case fit, Build OCR pipelines, Develop image attribute extraction systems, Work on image enhancement and generation, Create automated evaluation pipelines, Optimize ML systems across edge and cloud deployments

Seniority

Mid-Senior, hands-on IC

Rewrite
## About the Role ArcLogiq Softwares Pvt Ltd is hiring a Data Scientist / Applied ML Engineer to build and optimize DigiCatalog's probabilistic AI stack. The role focuses on selecting, orchestrating, and optimizing open-source AI models across OCR, multilingual NLP, computer vision, and multimodal pipelines. You'll work on scalable AI systems where cost, latency, and quality trade-offs are central to every technical decision. ## Responsibilities - Select and orchestrate open-source models based on use-case fit instead of defaulting to large LLMs. - Build OCR pipelines using: - DocTR - PaddleOCR - TrOCR-class models - Develop image attribute extraction systems using: - Grounding DINO - OWL-ViT - CLIP - Small VLMs (2–3B class) - Work on image enhancement and generation using: - Real-ESRGAN - SDXL-Turbo class models - Build Indic and multilingual AI workflows using: - IndicTrans2 - NLLB - IndicWhisper - Create automated evaluation pipelines and regression systems for: - OCR CER/WER - Attribute F1 - Hallucination rates - Language-specific performance slices - Perform error analysis, calibration, threshold tuning, ensembling, and post-processing. - Optimize ML systems across edge and cloud deployments for cost, latency, and quality. ## Requirements - 4–6 years of experience in Applied ML, CV, NLP, or Multimodal AI. - Strong hands-on expertise with OCR, CV, NLP, and open-source AI ecosystems. - Experience building production-grade ML pipelines and evaluation systems. - Strong Python and ML engineering skills. - Familiarity with inference optimization and model benchmarking. - Ability to quickly productionize open-source research and new models. ## Nice to Have - Experience with edge AI deployments. - Knowledge of vector search and embeddings. - Familiarity with GPU optimization and distributed inference pipelines. - Strong experimentation and research mindset. ## Hiring Process - Screening Round 1: Online assessment on Chaideep AI - Screening Round 2: Technical Interview - Screening Round 3: Final Discussion Managerial Round ## Why Join Us - Build cutting-edge multimodal AI systems. - Work on real-world AI deployment challenges at scale. - Ownership-driven and fast-paced engineering culture. - Opportunity to ship impactful AI products rapidly. ## Contact Interested candidates can share their updated resume at [email protected].
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗