PARSEME-AR: Arabic reference corpus for multiword expressions using PARSEME annotation guidelines
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3A8HXR5IS7" target="_blank" >RIV/00216208:11320/26:8HXR5IS7 - isvavai.cz</a>
Výsledek na webu
<a href="http://dx.doi.org/10.1007/s10579-024-09763-7" target="_blank" >http://dx.doi.org/10.1007/s10579-024-09763-7</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1007/s10579-024-09763-7" target="_blank" >10.1007/s10579-024-09763-7</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
PARSEME-AR: Arabic reference corpus for multiword expressions using PARSEME annotation guidelines
Popis výsledku v původním jazyce
In this paper we present PARSEME-AR, the first openly available Arabic corpus manually annotated for Verbal Multiword Expressions (VMWEs). The annotation process is carried out based on guidelines put forward by PARSEME, a multilingual project for more than 26 languages. The corpus contains 4749 VMWEs in about 7500 sentences taken from the Prague Arabic Dependency Treebank. The results notably show a high degree of discontinuity in Arabic VMWEs in comparison to other languages in the PARSEME suite. We also propose analyses of interesting and challenging phenomena encountered during the annotation process. Moreover, we offer the first benchmark for the VMWE identification task in Arabic, by training two state-of-the-art systems, on our Arabic data. © The Author(s), under exclusive licence to Springer Nature B.V. 2024.
Název v anglickém jazyce
PARSEME-AR: Arabic reference corpus for multiword expressions using PARSEME annotation guidelines
Popis výsledku anglicky
In this paper we present PARSEME-AR, the first openly available Arabic corpus manually annotated for Verbal Multiword Expressions (VMWEs). The annotation process is carried out based on guidelines put forward by PARSEME, a multilingual project for more than 26 languages. The corpus contains 4749 VMWEs in about 7500 sentences taken from the Prague Arabic Dependency Treebank. The results notably show a high degree of discontinuity in Arabic VMWEs in comparison to other languages in the PARSEME suite. We also propose analyses of interesting and challenging phenomena encountered during the annotation process. Moreover, we offer the first benchmark for the VMWE identification task in Arabic, by training two state-of-the-art systems, on our Arabic data. © The Author(s), under exclusive licence to Springer Nature B.V. 2024.
Klasifikace
Druh
J<sub>SC</sub> - Článek v periodiku v databázi SCOPUS
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
—
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
Language Resources and Evaluation
ISSN
1574-020X
e-ISSN
—
Svazek periodika
59
Číslo periodika v rámci svazku
2
Stát vydavatele periodika
US - Spojené státy americké
Počet stran výsledku
31
Strana od-do
1331-1361
Kód UT WoS článku
—
EID výsledku v databázi Scopus
2-s2.0-85202476399