Software module for automatic enhancement of digitized documents
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216305%3A26230%2F19%3APR32695" target="_blank" >RIV/00216305:26230/19:PR32695 - isvavai.cz</a>
Result on the web
<a href="https://www.fit.vut.cz/research/product/630/" target="_blank" >https://www.fit.vut.cz/research/product/630/</a>
DOI - Digital Object Identifier
—
Alternative languages
Result language
angličtina
Original language name
Software module for automatic enhancement of digitized documents
Original language description
Tool for text-guided textual document scan quality enhancement. The method works on lines of text that can be input through a PAGE XML or detected automatically by a built-in OCR. By using text input along with the image, the results can be correctly readable even with parts of the original text missing or severely degraded in the source image. The tool includes functionality for cropping the text lines, processing them with our provided models for either text enhancement and inpainting, and for blending the enhanced text lines back into the source document image. We currently provide models for OCR and enhancement of czech newspapers optimized for low-quality scans from micro-films. This package can be used as a standalone command line tool to process document pages in bulk. Alternatively, the package provides a python class that can be integrated in third-party software.
Czech name
—
Czech description
—
Classification
Type
R - Software
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
<a href="/en/project/DG18P02OVV055" target="_blank" >DG18P02OVV055: Advanced content extraction and recognition for printed and handwritten documents for better accessibility and usability</a><br>
Continuities
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Others
Publication year
2019
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Internal product ID
PERO-ENHANCE
Technical parameters
Využití na základě volné a bezplatné open-source licence.
Economical parameters
Jedná se o modul pro integraci do digitalizačních linek a digitalizačního software. Komerční uplatnění je možné v rámci poskytování doplňkových služeb a konzultací.
Owner IČO
—
Owner name
Fakulta informačních technologií