Phrasemes and Collocations in the Corpus – How to Find Unknown Variants
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11210%2F25%3A10507558" target="_blank" >RIV/00216208:11210/25:10507558 - isvavai.cz</a>
Nalezeny alternativní kódy
RIV/00216208:11320/26:EUPBQ29Q
Výsledek na webu
<a href="https://verso.is.cuni.cz/pub/verso.fpl?fname=obd_publikace_handle&handle=0tNO2RvpqV" target="_blank" >https://verso.is.cuni.cz/pub/verso.fpl?fname=obd_publikace_handle&handle=0tNO2RvpqV</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.2478/jazcas-2025-0019" target="_blank" >10.2478/jazcas-2025-0019</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Phrasemes and Collocations in the Corpus – How to Find Unknown Variants
Popis výsledku v původním jazyce
This paper addresses the identification and annotation of multiwordexpressions (MWEs) in Czech corpora, focusing on enhancing the search procedurethrough transformations of existing lexicon entries and the addition of new entries based onsyntactic patterns. We discuss the limitations of current annotation systems and introducea new, efficient annotation system that leverages a comprehensive MWE dictionary. Ourmethodology includes the use of syntactic patterns to identify new collocations, automatictransformations of known MWEs, and manual searches for creatively varied expressions.The results demonstrate significant improvements in the success rate of corpus annotation,with newly identified collocations and transformed MWEs contributing to a richer and moreaccurate linguistic resource.
Název v anglickém jazyce
Phrasemes and Collocations in the Corpus – How to Find Unknown Variants
Popis výsledku anglicky
This paper addresses the identification and annotation of multiwordexpressions (MWEs) in Czech corpora, focusing on enhancing the search procedurethrough transformations of existing lexicon entries and the addition of new entries based onsyntactic patterns. We discuss the limitations of current annotation systems and introducea new, efficient annotation system that leverages a comprehensive MWE dictionary. Ourmethodology includes the use of syntactic patterns to identify new collocations, automatictransformations of known MWEs, and manual searches for creatively varied expressions.The results demonstrate significant improvements in the success rate of corpus annotation,with newly identified collocations and transformed MWEs contributing to a richer and moreaccurate linguistic resource.
Klasifikace
Druh
J<sub>SC</sub> - Článek v periodiku v databázi SCOPUS
CEP obor
—
OECD FORD obor
60203 - Linguistics
Návaznosti výsledku
Projekt
Výsledek vznikl pri realizaci vícero projektů. Více informací v záložce Projekty.
Návaznosti
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
Jazykovedný Časopis
ISSN
0021-5597
e-ISSN
1338-4287
Svazek periodika
76
Číslo periodika v rámci svazku
1
Stát vydavatele periodika
SK - Slovenská republika
Počet stran výsledku
11
Strana od-do
212-222
Kód UT WoS článku
—
EID výsledku v databázi Scopus
2-s2.0-105025763252