Aranea Go Middle East: Persicum
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F22%3AMG8E7ZQC" target="_blank" >RIV/00216208:11320/22:MG8E7ZQC - isvavai.cz</a>
Result on the web
<a href="https://nlp.fi.muni.cz/raslan/raslan22.pdf#page=113" target="_blank" >https://nlp.fi.muni.cz/raslan/raslan22.pdf#page=113</a>
DOI - Digital Object Identifier
—
Alternative languages
Result language
angličtina
Original language name
Aranea Go Middle East: Persicum
Original language description
Our paper introduces the creation and annotation of Araneum Persicum, a new Persian web-crawled corpus. Some problems encountered during the process of filtration and annotation are shown, and an ensemble approach adopted for lemmatization and morphosyntactic annotation is introduced. It is also argued that Romanization can be helpful in developing corpora for languages not based on Latin script.
Czech name
—
Czech description
—
Classification
Type
D - Article in proceedings
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
—
Continuities
—
Others
Publication year
2022
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Article name in the collection
RASLAN 2022 Recent Advances in Slavonic Natural Language Processing
ISBN
978-80-263-1752-4
ISSN
—
e-ISSN
—
Number of pages
9
Pages from-to
113-121
Publisher name
Tribun EU
Place of publication
—
Event location
Karlova Studánka, Czech Republic
Event date
Jan 1, 2022
Type of event by nationality
WRD - Celosvětová akce
UT code for WoS article
—