An Automated Pipeline for Robust Image Processing and Optical Character Recognition of Historical Documents
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F49777513%3A23520%2F20%3A43959664" target="_blank" >RIV/49777513:23520/20:43959664 - isvavai.cz</a>
Result on the web
<a href="https://link.springer.com/chapter/10.1007/978-3-030-60276-5_17" target="_blank" >https://link.springer.com/chapter/10.1007/978-3-030-60276-5_17</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1007/978-3-030-60276-5_17" target="_blank" >10.1007/978-3-030-60276-5_17</a>
Alternative languages
Result language
angličtina
Original language name
An Automated Pipeline for Robust Image Processing and Optical Character Recognition of Historical Documents
Original language description
In this paper, we propose a pipeline for processing of scanned historical documents into the electronic text form that could then be indexed and stored in a database. The nature of the documents presents a substantial challenge for standard automated techniques — not only there is a mix of typewritten and handwritten documents of varying quality but the scanned pages often contain multiple documents at once. Moreover, the language of the texts alternates mostly between Russian and Ukrainian but other languages also occur. The paper focuses mainly on segmentation, document type classification, and image preprocessing of the scanned documents; the output of those methods is then passed to the off-the-shelf OCR software and a baseline performance is evaluated on a simplified OCR task.
Czech name
—
Czech description
—
Classification
Type
D - Article in proceedings
CEP classification
—
OECD FORD branch
20205 - Automation and control systems
Result continuities
Project
<a href="/en/project/DG20P02OVV018" target="_blank" >DG20P02OVV018: Digital archive of the NKVD/KGB files related to Czechoslovakia</a><br>
Continuities
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Others
Publication year
2020
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Article name in the collection
Speech and Computer, 22nd International Conference, SPECOM 2019, St. Petersburg, Russia, October 7-9,2020, Proceedings
ISBN
978-3-030-60275-8
ISSN
0302-9743
e-ISSN
1611-3349
Number of pages
10
Pages from-to
166-175
Publisher name
Springer
Place of publication
Cham
Event location
St. Petersburg, Russia
Event date
Oct 7, 2020
Type of event by nationality
WRD - Celosvětová akce
UT code for WoS article
—