Lemmatization of Czech and Croatian Noun Clusters for Terminology Extraction
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216224%3A14330%2F25%3A00142890" target="_blank" >RIV/00216224:14330/25:00142890 - isvavai.cz</a>
Výsledek na webu
<a href="https://nlp.fi.muni.cz/raslan/2025/" target="_blank" >https://nlp.fi.muni.cz/raslan/2025/</a>
DOI - Digital Object Identifier
—
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Lemmatization of Czech and Croatian Noun Clusters for Terminology Extraction
Popis výsledku v původním jazyce
During terminology extraction, terms discovered in corpora are presented in their canonical form. Lemmatization of multi-word terms consisting of noun clusters can be ambiguous due to the lack of information on their internal structure. In this paper, we show that grammatical case alone is often not sufficient for the construction of canonical forms of noun clusters. We focus on two-noun clusters in the genitive, which are the most frequent type with ambiguous parsing. Based on corpus research, we design rules that make use of multiple morphological categories to improve the lemmatization of noun clusters found in Czech and Croatian corpora. In addition to case, we also take note of gender, animacy, and whether the noun is a proper noun. The improvements lead to more accurate and more unified forms of the terms produced during terminology extraction for these two languages in Sketch Engine.
Název v anglickém jazyce
Lemmatization of Czech and Croatian Noun Clusters for Terminology Extraction
Popis výsledku anglicky
During terminology extraction, terms discovered in corpora are presented in their canonical form. Lemmatization of multi-word terms consisting of noun clusters can be ambiguous due to the lack of information on their internal structure. In this paper, we show that grammatical case alone is often not sufficient for the construction of canonical forms of noun clusters. We focus on two-noun clusters in the genitive, which are the most frequent type with ambiguous parsing. Based on corpus research, we design rules that make use of multiple morphological categories to improve the lemmatization of noun clusters found in Czech and Croatian corpora. In addition to case, we also take note of gender, animacy, and whether the noun is a proper noun. The improvements lead to more accurate and more unified forms of the terms produced during terminology extraction for these two languages in Sketch Engine.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10200 - Computer and information sciences
Návaznosti výsledku
Projekt
<a href="/cs/project/LM2023062" target="_blank" >LM2023062: Digitální výzkumná infrastruktura pro jazykové technologie, umění a humanitní vědy</a><br>
Návaznosti
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
Recent Advances in Slavonic Natural Language Processing, RASLAN 2025
ISBN
9788026318583
ISSN
2336-4289
e-ISSN
—
Počet stran výsledku
8
Strana od-do
165-172
Název nakladatele
Tribun EU
Místo vydání
Brno, Czech Republic
Místo konání akce
Kouty nad Desnou, Česká Republika
Datum konání akce
1. 1. 2025
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
—