Multi-label Classification and Named Entity Recognition for Historical Documents
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F49777513%3A23520%2F25%3A43973196" target="_blank" >RIV/49777513:23520/25:43973196 - isvavai.cz</a>
Výsledek na webu
<a href="https://link.springer.com/chapter/10.1007/978-3-031-81010-7_2" target="_blank" >https://link.springer.com/chapter/10.1007/978-3-031-81010-7_2</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1007/978-3-031-81010-7_2" target="_blank" >10.1007/978-3-031-81010-7_2</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Multi-label Classification and Named Entity Recognition for Historical Documents
Popis výsledku v původním jazyce
In this paper, we present improvements to our processing pipeline for historical document digitization. The original pipeline is extended with two new functionalities - page labeling, and named entity recognition. We handle page labeling as a multi-label classification task, for which we choose the Query2Label approach. Query2Label is tested on our internal NKVD dataset and reaches a mean average precision equal to 80.03% on the test set. For the named entity recognition task we utilize pre-trained transformer-based models DeepPavlov and benchmark them on two entities - person name, and location. The best model reaches promising results despite not being trained on our data at all.
Název v anglickém jazyce
Multi-label Classification and Named Entity Recognition for Historical Documents
Popis výsledku anglicky
In this paper, we present improvements to our processing pipeline for historical document digitization. The original pipeline is extended with two new functionalities - page labeling, and named entity recognition. We handle page labeling as a multi-label classification task, for which we choose the Query2Label approach. Query2Label is tested on our internal NKVD dataset and reaches a mean average precision equal to 80.03% on the test set. For the named entity recognition task we utilize pre-trained transformer-based models DeepPavlov and benchmark them on two entities - person name, and location. The best model reaches promising results despite not being trained on our data at all.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
20205 - Automation and control systems
Návaznosti výsledku
Projekt
<a href="/cs/project/DH23P03OVV073" target="_blank" >DH23P03OVV073: Databáze pramenů k problematice politických represí vůči čs. občanům a krajanům v Sovětském svazu</a><br>
Návaznosti
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
Dynamics of Information Systems. DIS 2024. Lecture Notes in Computer Science
ISBN
978-3-031-81009-1
ISSN
0302-9743
e-ISSN
1611-3349
Počet stran výsledku
11
Strana od-do
24-34
Název nakladatele
Springer
Místo vydání
Cham
Místo konání akce
Kalamata, Greece
Datum konání akce
2. 6. 2024
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
001534826000002