Automatic document quality assessment software module
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216305%3A26230%2F19%3APR32697" target="_blank" >RIV/00216305:26230/19:PR32697 - isvavai.cz</a>
Result on the web
<a href="https://github.com/DCGM/pero-quality" target="_blank" >https://github.com/DCGM/pero-quality</a>
DOI - Digital Object Identifier
—
Alternative languages
Result language
angličtina
Original language name
Automatic document quality assessment software module
Original language description
This tool provides automatic quality assessment of digitalized documents. The estimated quality scores closely correspond to readability by humans. The tool provides quality score heatmaps and an overall quality score for a whole document page. The module computes local perceptual quality scores based on confidence scores from Optical Character Recognition (OCR) or directly by a fast convolutional neural network. This module is build on top of OCR developed in project PERO (pero-ocr). The text recognition works in multiple stages. Firstly, locations and heights of text lines are determined using a fully convolutional neural network (modified U-NET). The individual text lines are processed by covolutional-recurrent networks trained using CTC loss. These networks provide confidences of recognized characters which are locally mapped to perceptual scores. The mapping to perceptual scores was calibrated on a large dataset of readability ratings by human readers.
Czech name
—
Czech description
—
Classification
Type
R - Software
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
<a href="/en/project/DG18P02OVV055" target="_blank" >DG18P02OVV055: Advanced content extraction and recognition for printed and handwritten documents for better accessibility and usability</a><br>
Continuities
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Others
Publication year
2019
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Internal product ID
PERO-QUALITY
Technical parameters
Využití na základě volné a bezplatné open-source licence.
Economical parameters
Jedná se o modul pro integraci do digitalizačních linek a digitalizačního software. Komerční uplatnění je možné v rámci poskytování doplňkových služeb a konzultací.
Owner IČO
—
Owner name
Fakulta informačních technologií