CV / 多模态算法实习生(J105451)
Core
Develop algorithms for document scanning and conversion in real-world scenarios, including image-to-Word/Excel conversion and document understanding.
Role type
Computer Vision / Multimodal Algorithm Intern
Builds
Document scanning and document conversion features
Domain
Computer Vision, OCR, Document Understanding
Deliverable
production ML models
Required skills
Computer Vision, Deep Learning, Image Classification, Object Detection, Image Segmentation, OCR, Layout Analysis, Python, PyTorch
Preferred skills
Paper reading and reproduction, applying cutting-edge CV/VLM methods to real problems
Technologies
PyTorch
Responsibilities
Develop algorithms for document region detection, perspective correction, image enhancement, and noise removal; Optimize algorithm accuracy and inference efficiency by analyzing bad cases; Collaborate with product and engineering teams to deploy algorithms.
Seniority
Intern