Python Backend Developer
Core
Build a Python-based pipeline to convert complex, unstructured PDFs into perfectly formatted, editable .docx files by integrating OCR engines and AI models.
Role type
Senior IC backend developer (document processing & AI)
Builds
Document conversion pipeline for translation software
Domain
AI / Machine Translation / Document Processing
Deliverable
production ML models | product features
Required skills
Python, OpenCV, PyMuPDF, python-docx, OCR (Tesseract, PaddleOCR), Transformer models (LayoutLMv3, Donut, Nougat), LLM integration (GPT, Claude), PyTorch, TensorFlow
Preferred skills
Pandoc AST, DTP/Typography knowledge, Open-source OCR contributions
Technologies
Python, OpenCV, PyMuPDF, python-docx, Tesseract, PaddleOCR, LayoutLMv3, Donut, Nougat, GPT, Claude, PyTorch, TensorFlow
Responsibilities
Perform comparative analysis of commercial vs. open-source OCR solutions; Establish metrics for format fidelity; Build Python workflow integrating OCR with document generation libraries; Explore and fine-tune Vision-Language Models for structural recognition; Solve edge cases like rotated text and complex notation
Seniority
Senior, hands-on IC