Czech Historical Named Entity Corpus v 1.0
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F49777513%3A23520%2F20%3A43961403" target="_blank" >RIV/49777513:23520/20:43961403 - isvavai.cz</a>
Result on the web
<a href="https://www.aclweb.org/anthology/2020.lrec-1.549.pdf" target="_blank" >https://www.aclweb.org/anthology/2020.lrec-1.549.pdf</a>
DOI - Digital Object Identifier
—
Alternative languages
Result language
angličtina
Original language name
Czech Historical Named Entity Corpus v 1.0
Original language description
As the number of digitized archival documents increases very rapidly, named entity recognition (NER) in historical documents has become very important for information extraction and data mining. For this task an annotated corpus is needed, which has up to now been missing for Czech. In this paper we present a new annotated data collection for historical NER, composed of Czech historical newspapers. This corpus is freely available for research purposes at http://chnec.kiv.zcu.cz/. For this corpus, we have defined relevant domain-specific named entity types and created an annotation manual for corpus labelling. We further conducted some experiments on this corpus using recurrent neural networks in order to show baseline results on this dataset. We experimented with randomly initialized embeddings and static and dynamic fastText word embeddings. We achieved 0.73 F1 score with a bidirectional LSTM model using static fastText embeddings.
Czech name
—
Czech description
—
Classification
Type
D - Article in proceedings
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
—
Continuities
O - Projekt operacniho programu
Others
Publication year
2020
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Article name in the collection
Proceedings of the 12th Language Resources and Evaluation Conference
ISBN
979-10-95546-34-4
ISSN
—
e-ISSN
—
Number of pages
8
Pages from-to
4458-4465
Publisher name
European Language Resources Association (ELRA)
Place of publication
Paris
Event location
Marseille, France
Event date
May 11, 2020
Type of event by nationality
WRD - Celosvětová akce
UT code for WoS article
—